Reinforcement learning-based method and system for site selection optimization of chain stores

By constructing a reinforcement learning system to monitor urban planning and traffic information in real time, and by using a three-dimensional digital twin environment and an adaptive weighted reward function to optimize site selection strategies, the problem of inaccurate store site selection in existing technologies has been solved, realizing dynamic and intelligent site selection decisions and improving site selection effectiveness.

CN121860694BActive Publication Date: 2026-06-23ZHEJIANG GUMING HOUAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG GUMING HOUAN INFORMATION TECH CO LTD
Filing Date
2026-03-17
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing methods for selecting store locations in the chain retail industry cannot respond in real time to the dynamic changes in urban planning and transportation networks, resulting in inaccurate location selection, an inability to dynamically balance rental costs, expected customer traffic and return on investment, and a lack of an iterative optimization mechanism for decision-making strategies.

Method used

We construct a store location selection system based on reinforcement learning. By monitoring urban planning and traffic information in real time, using a three-dimensional digital twin environment and an adaptive weighted reward function, we dynamically adjust decision weights and conduct periodic training based on actual operational data to optimize the location selection strategy.

Benefits of technology

It enables dynamic, precise, and intelligent site selection for chain stores, improves the scientific and rational nature of site selection decisions, helps chain brands accurately target high-potential locations, and enhances store operating efficiency and market competitiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860694B_ABST
    Figure CN121860694B_ABST
Patent Text Reader

Abstract

The application provides a chain store site selection optimization method and system based on reinforcement learning, and relates to the technical field of artificial intelligence. The method comprises the following steps: constructing a reinforcement learning environment for store site selection, initializing a reinforcement learning intelligent agent, and monitoring the change signals of external city planning and traffic information in real time; based on the monitored change signals, selecting three bus hubs as evaluation base points in the city geographical space, and defining the core area of the evaluated business circle by the three bus hub base points. The application realizes the intelligentization, dynamization and precision of site selection decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method and system for optimizing the location of chain stores based on reinforcement learning. Background Technology

[0002] Against the backdrop of rapid expansion in the chain retail industry, the scientific nature of store site selection directly determines the success or failure of operations. Traditional site selection methods mostly rely on human experience judgment or static data statistical analysis, which makes it difficult to accurately cope with dynamic environmental changes such as urban planning adjustments, transportation network optimization, and population flow changes. Although some existing site selection technologies have attempted to incorporate data mining methods, they lack the ability to adapt to dynamic environments in real time and the iterative optimization mechanism for decision-making strategies.

[0003] When a chain convenience store brand was setting up new stores in an emerging urban area, it used a site selection scheme based on the distribution of historical commercial facilities and static population data to determine the store locations. However, it failed to capture the reconstruction of customer flow distribution brought about by the subsequent relocation of public transportation hubs in the area, and could not dynamically balance the relationship between rental costs, expected customer flow and return on investment. As a result, the new stores had insufficient customer traffic and lower-than-expected operating efficiency. This case highlights the shortcomings of existing site selection technologies in terms of dynamic environmental perception, multi-objective decision-making balance, and strategy iteration optimization. Summary of the Invention

[0004] The technical problem to be solved by this invention is to provide a method and system for optimizing the site selection of chain stores based on reinforcement learning, so as to realize the intelligent, dynamic and precise site selection decision.

[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:

[0006] Firstly, a chain store site selection optimization method based on reinforcement learning, the method comprising:

[0007] Construct a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time;

[0008] Based on the monitored change signals, three public transport hubs were selected in the urban geospatial space as evaluation base points, and the core area of ​​the business district to be evaluated was defined by the three public transport hub base points.

[0009] By dividing the core area of ​​the business district into grids, the geometric moments and central moments of the population distribution point set and commercial facility point set within each grid unit are used to quantify the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, the population density, commercial popularity and transportation convenience indicators within each grid are comprehensively evaluated to obtain regional characteristic compensation parameters.

[0010] The original monitoring data is calibrated using regional feature compensation parameters to obtain corrected monitoring results; the reinforcement learning environment is updated based on the monitoring results to obtain the current environmental state.

[0011] The current environmental state is input into the reinforcement learning agent; the agent dynamically adjusts the weights of rental costs, expected customer traffic and return on investment in the reward calculation by calling the adaptive weighted reward function, and evaluates the candidate store locations in combination with the environmental state to obtain the site selection strategy.

[0012] By opening new stores through site selection strategies, actual operational return data is collected. This operational return data is used as environmental feedback, which, together with the corresponding environmental conditions and site selection strategies, constitutes training samples for periodic training of the reinforcement learning agent to optimize decision-making strategies.

[0013] Furthermore, a reinforcement learning environment for store site selection is constructed, and a reinforcement learning agent is initialized to monitor changes in external urban planning and traffic information in real time, including:

[0014] Urban planning information and traffic network data are obtained from urban planning databases, data interfaces of traffic management departments, and third-party geographic information to form the original monitoring dataset;

[0015] By cleaning, standardizing the format, and aligning the spatial coordinates of the original monitoring dataset, structured urban spatial basic data is obtained.

[0016] A three-dimensional digital twin environment, including geospatial grids, population flow trajectories, and the distribution of commercial facilities, is constructed based on structured urban spatial data as a foundational environmental framework for reinforcement learning.

[0017] The initial parameters of the reinforcement learning agent are configured based on the three-dimensional digital twin environment, including the state space dimension, action space range and initial policy network weights, to complete the initial deployment of the agent;

[0018] Based on the reinforcement learning agent and 3D digital twin environment deployed in the initial stage, a real-time data connection channel is established with the urban planning department and traffic management platform. The data update frequency and change detection threshold are set to obtain a dynamic data flow input mechanism.

[0019] Based on a dynamic data stream input mechanism, the system continuously receives updated urban planning and traffic information, integrates the updated data with the three-dimensional digital twin environment in real time, and generates environmental status change signals.

[0020] Furthermore, based on the monitored change signals, three public transport hubs were selected as evaluation benchmarks in the urban geographic space. These three public transport hub benchmarks jointly define the core area of ​​the business district to be evaluated, including:

[0021] Based on signals of changes in environmental conditions, we analyze the degree of impact of changes in urban planning and transportation network adjustments, and identify the urban areas affected by these changes.

[0022] Based on the urban area affected by the changes, data on all public transport hub nodes within the area are extracted from the three-dimensional digital twin environment. Combined with historical passenger flow data, number of transfer lines, and population density indicators of the hubs, a candidate public transport hub evaluation list is obtained.

[0023] Based on the candidate bus hub evaluation list, the spatial distribution balance algorithm is used to calculate the geometric distance and passenger flow complementarity coefficient between each candidate hub, and the three bus hubs with the most balanced spatial distribution and the highest passenger flow coverage complementarity are selected as evaluation base points.

[0024] By assessing the spatial coordinates of three public transport hubs as evaluation base points, a triangular coverage area is constructed. By calculating the radius of the circumcircle connecting the three base points and the regional population density decay function, the boundary range of the core business district is determined, thus obtaining the core business district to be evaluated.

[0025] Furthermore, by dividing the core area of ​​the business district into grids, and based on the geometric moments and central moments of the population distribution point set and commercial facility point set within each grid unit, the distribution concentration and dispersion balance characteristics of various point sets are quantified. Simultaneously, combined with the moment characteristics of the transportation station distribution point set within the grid, the population density, commercial activity, and transportation convenience indicators within each grid are comprehensively evaluated to obtain regional characteristic compensation parameters, including:

[0026] The core area of ​​the business district is divided into regular grids according to a preset resolution to obtain a grid set. The point set data of population distribution, commercial facilities and transportation stations in each grid are extracted from the three-dimensional digital twin environment. The geometric moments and central moments of the population and commercial facility point sets are calculated by the shape moment calculation algorithm to obtain the population distribution concentration parameter and the commercial facility distribution balance parameter.

[0027] The moment characteristics of the traffic station distribution point set within each grid are calculated using the population distribution concentration parameter and the commercial facility distribution balance parameter, including the station distribution compactness parameter and coverage uniformity parameter.

[0028] Based on concentration parameters, balance parameters, and traffic station moment characteristics, a multi-dimensional fusion evaluation method is used to comprehensively calculate the population density index, commercial popularity index, and transportation convenience index of each grid, resulting in a grid comprehensive evaluation vector.

[0029] Based on the grid-based comprehensive evaluation vector, the regional feature compensation parameter matrix for calibrating the original monitoring data is obtained through regional feature difference analysis and standardization processing.

[0030] Furthermore, the original monitoring data is calibrated using regional feature compensation parameters to obtain corrected monitoring results; the reinforcement learning environment is updated based on the monitoring results to obtain the current environmental state, including:

[0031] The structured urban spatial baseline data is calibrated grid by grid based on the regional feature compensation parameter matrix to obtain calibrated spatial data;

[0032] By performing data fusion and smoothing on the calibrated spatial data, the corrected monitoring results are obtained;

[0033] The constructed three-dimensional digital twin environment is dynamically updated based on the corrected monitoring results to obtain updated environmental data;

[0034] The current environmental state is obtained by extracting key state features from the updated environmental data.

[0035] Furthermore, the current environmental state is input into the reinforcement learning agent; the agent dynamically adjusts the weights of rental costs, expected customer traffic, and return on investment in the reward calculation by calling an adaptive weighted reward function, and evaluates candidate locations in conjunction with the environmental state to obtain a site selection strategy, including:

[0036] By receiving the current environmental state, the current environmental state is input into the initialized reinforcement learning agent; based on the current environmental state, the adaptive weighted reward function built into the reinforcement learning agent is called, and the weight ratio of the three indicators of rental cost, expected passenger flow and investment return rate in the reward calculation is dynamically adjusted according to the changing trend of urban economic activity and historical operating data.

[0037] Based on dynamically adjusted weight ratios and the current environmental status, a multi-dimensional evaluation is conducted on all candidate geographic grid locations within the core area of ​​the business district, and a comprehensive reward score is calculated for each candidate location.

[0038] Candidate locations are ranked and screened based on comprehensive reward scores. Combined with regional competitive landscape analysis and brand positioning requirements, the final site selection strategy is obtained, including recommended store location coordinates and expected operating efficiency indicators.

[0039] Furthermore, by opening new stores through site selection strategies, actual operational return data is collected. This operational return data, along with the corresponding environmental state and site selection strategy, forms training samples for periodic training of the reinforcement learning agent to optimize decision-making strategies, including:

[0040] By receiving site selection strategies, instructions for opening new stores and recommended store location coordinates are sent to store management.

[0041] Based on the execution results of the new store opening instructions, the actual operational return data of the new stores is collected in real time, including average daily customer traffic, single store revenue and cost expenditure data;

[0042] Based on actual operational return data, operational performance indicators are calculated and used as environmental feedback signals.

[0043] Based on environmental feedback signals, the corresponding environmental states and location strategies are associated to construct a complete training sample dataset;

[0044] Using the training sample dataset, the parameters of the initialized reinforcement learning agent are updated and trained according to a preset time period.

[0045] The optimized decision-making strategy is obtained by updating the parameters of the trained reinforcement learning agent.

[0046] Secondly, a chain store site selection optimization system based on reinforcement learning includes:

[0047] The monitoring module is used to build a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time.

[0048] The delineation module is used to select three public transport hubs as evaluation base points in the urban geospatial space based on the monitored change signals, and to jointly delineate the core area of ​​the business district to be evaluated using the three public transport hub base points.

[0049] The partitioning module is used to divide the core area of ​​the business district into grids. Based on the geometric moments and central moments of the population distribution point set and the commercial facility point set within each grid unit, it quantifies the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, it comprehensively evaluates the population density, commercial popularity and transportation convenience indicators within each grid to obtain regional characteristic compensation parameters.

[0050] The correction module is used to calibrate the original monitoring data using regional feature compensation parameters to obtain corrected monitoring results; and to update the reinforcement learning environment based on the monitoring results to obtain the current environmental state.

[0051] The calculation module is used to input the current environmental state into the reinforcement learning agent; the agent dynamically adjusts the weights of rental cost, expected customer flow and return on investment in the reward calculation by calling the adaptive weighted reward function, and evaluates the candidate store sites in combination with the environmental state to obtain the site selection strategy.

[0052] The processing module is used to open new stores through site selection strategies and collect actual operational return data. The operational return data is used as environmental feedback, which, together with the corresponding environmental state and site selection strategy, constitutes training samples to periodically train the reinforcement learning agent in order to optimize the decision-making strategy.

[0053] Thirdly, a computing device includes:

[0054] One or more processors;

[0055] A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.

[0056] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.

[0057] The above-described solution of the present invention has at least the following beneficial effects:

[0058] By employing techniques to construct a 3D digital twin environment and establish real-time data connections and dynamic data stream input mechanisms, this technology overcomes the technical problem that existing site selection technologies cannot respond in real time to the dynamic changes in urban planning and transportation networks. By using shape moment calculation algorithms to quantify regional spatial characteristics and generate regional feature compensation parameter matrices, it overcomes the technical problem of data bias caused by inaccurate quantification of regional population, commercial, and transportation characteristics in existing technologies. Furthermore, by employing adaptive weighted reward functions to dynamically adjust multi-objective decision weights and constructing a reinforcement learning closed-loop training mechanism based on actual operational data, it overcomes the technical problem that existing technologies struggle to balance the multi-objective demands of rental costs, expected customer flow, and return on investment, and lack the ability to iteratively optimize decision-making strategies. This achieves dynamic, precise, and intelligent site selection decisions for chain stores, improving the scientific and rational nature of site selection decisions, helping chain brands accurately target high-potential locations, and enhancing store operating efficiency and market competitiveness. Attached Figure Description

[0059] Figure 1 This is a flowchart illustrating the chain store site selection optimization method based on reinforcement learning provided in an embodiment of the present invention.

[0060] Figure 2 This is a schematic diagram of a chain store site selection optimization system based on reinforcement learning provided in an embodiment of the present invention. Detailed Implementation

[0061] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.

[0062] like Figure 1 As shown, embodiments of the present invention propose a chain store site selection optimization method based on reinforcement learning, the method comprising the following steps:

[0063] Step 1: Build a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time.

[0064] Step 2: Based on the monitored change signals, select three public transport hubs in the urban geographic space as evaluation base points, and use the three public transport hub base points to jointly define the core area of ​​the business district to be evaluated;

[0065] Step 3: By dividing the core area of ​​the business district into grids, the geometric moments and central moments of the population distribution point set and commercial facility point set within each grid unit are used to quantify the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, the population density, commercial popularity and transportation convenience indicators within each grid are comprehensively evaluated to obtain regional characteristic compensation parameters.

[0066] Step 4: The original monitoring data is calibrated using regional feature compensation parameters to obtain the corrected monitoring results; the reinforcement learning environment is updated based on the monitoring results to obtain the current environmental state.

[0067] Step 5: Input the current environmental state into the reinforcement learning agent; the agent dynamically adjusts the weights of rental costs, expected customer traffic and return on investment in the reward calculation by calling the adaptive weighted reward function, and evaluates the candidate store locations in combination with the environmental state to obtain the site selection strategy.

[0068] Step 6: Open new stores through site selection strategies and collect actual operational return data; use the operational return data as environmental feedback, and together with the corresponding environmental state and site selection strategy, form training samples to periodically train the reinforcement learning agent in order to optimize the decision-making strategy.

[0069] In this embodiment of the invention, by employing techniques such as constructing a reinforcement learning environment, initializing an intelligent agent, and monitoring changes in urban planning and traffic information in real time, combined with the selection of public transport hubs and the definition of core business districts, grid division, and the quantification of regional features using shape moment algorithms to obtain compensation parameters, and then updating the environmental state through parameter calibration, the intelligent agent calls an adaptive weighted reward function to evaluate candidate store locations, and simultaneously uses actual operational data to construct training samples for periodic training and optimization of the intelligent agent, the technical problems of traditional site selection methods, such as inability to respond to dynamic changes in the urban environment in real time, inaccurate quantification of regional features, imbalance of multi-objective decision weights, and lack of iterative optimization mechanisms for decision-making strategies, are overcome. This achieves dynamic adaptation, accurate evaluation, and continuous optimization of chain store site selection decisions, improves the scientificity and rationality of site selection decisions, helps chain brands accurately lock in high-potential store locations, and enhances store operating efficiency and market competitiveness.

[0070] In a preferred embodiment of the present invention, step 1 above may include:

[0071] Step 1.1: Obtain urban planning information and transportation network data from urban planning databases, data interfaces of transportation management departments, and third-party geographic information sources to form the original monitoring dataset. Specifically, this includes: obtaining complete urban planning information from the official databases of urban planning authorities, including relevant documents such as the city master plan, zoning plan, control detailed plan, and special plan, covering core indicator data such as land use, plot ratio, building density, green space ratio, and urban infrastructure supporting plans; obtaining comprehensive transportation network data from standardized data interfaces opened by transportation management departments, including information on road grades, directions, lengths, number of lanes, and traffic capacity of urban road networks; specific directions, station settings, operating hours, and departure intervals of bus and subway lines; real-time traffic flow monitoring data; traffic accident statistics; and traffic control information. Finally, obtaining high-precision geographic information data from legally qualified third-party geographic information system service providers, including urban electronic maps, topographic data, administrative boundary data, and POI (Point of Interest) data, where POI data covers various spatial locations related to store location selection, such as commercial facilities, transportation stations, residential areas, schools, and hospitals. All the data obtained from the three different channels were categorized and summarized into three major categories: urban planning information, transportation network data, and geographic information data, forming a raw monitoring dataset that comprehensively covers the basic information required for store site selection.

[0072] Step 1.2 involves cleaning, standardizing, and aligning the original monitoring dataset to obtain structured urban spatial foundation data. Specifically, this includes: conducting comprehensive data cleaning of the original monitoring dataset; for data with missing values, if the missing percentage is less than 3%, filling is done using the mean or median of the field; if the missing percentage is greater than 3%, reasonable estimation and supplementation are performed based on the data's category and the characteristics of related data; for outliers, by comparing the distribution range of the data with data of the same category, data exceeding three times the standard deviation of the mean are identified as outliers and replaced with normal data from adjacent time periods or regions; for duplicate data, by comparing the core field information row by row, completely duplicate or... Redundant data with consistent core information ensures the uniqueness of the dataset. Next, format standardization is performed, converting urban planning information from different sources into Extensible Markup Language (XML) format, transportation network data into the industry-standard vector data format for Geographic Information Systems (GIS), and geographic information data into raster data format, ensuring all data has a unified storage format and access standards. Finally, spatial coordinate alignment is performed, converting all datasets involving spatial location into the National Geodetic Coordinate System 2000. Coordinate transformation algorithms are used to accurately calculate the coordinates of each spatial point, eliminating positional deviations caused by different coordinate systems and ensuring all data maintains spatial consistency. This ultimately results in structured, accurate, and spatially aligned urban spatial foundation data.

[0073] Step 1.3 involves constructing a 3D digital twin environment based on structured urban spatial data, including geospatial grids, population flow trajectories, and commercial facility distribution. This serves as the foundational framework for reinforcement learning. Specifically, this includes: First, based on the structured urban spatial data, a geospatial grid is divided. Using a fixed resolution of 100 meters by 100 meters, and with the city's administrative boundaries as the boundary, the urban space is divided into several regular square grids. Each grid is assigned a unique identifier, and the spatial boundary coordinates of each grid are clearly defined. Then, population flow trajectory data is collected. By integrating anonymized mobile phone signaling location data provided by telecommunications operators, bus and subway card swiping data from public transportation operators, and shared bicycle riding trajectory data, key information such as the origin, destination, route, and dwell time of different groups of people at different time periods is extracted to construct a population flow trajectory dataset reflecting the real-time flow patterns of the urban population. Simultaneously, commercial facility distribution data is organized. All commercially relevant Points of Interest (POIs) are selected from the structured urban spatial data, including existing chain stores, shopping malls, supermarkets, restaurants, and entertainment facilities. Detailed information such as the specific location, business type, operating area, and operating hours of each commercial facility is clearly defined. By deeply integrating the divided geospatial grid, the constructed population flow trajectory dataset, and the organized commercial facility distribution data, and using digital twin modeling technology, a 3D virtual scene containing elements such as urban terrain, buildings, road networks, commercial facilities, and population flow simulation is constructed through 3D modeling software. The 3D virtual scene realistically restores the physical characteristics and dynamic changes of urban space, serving as the basic environmental framework for reinforcement learning.

[0074] Step 1.4: Configure the initial parameters of the reinforcement learning agent based on the 3D digital twin environment, including the state space dimension, action space range, and initial policy network weights, to complete the initial deployment of the agent. Specifically, this includes: determining the state space dimension of the reinforcement learning agent based on the constructed 3D digital twin environment; the state space dimension covers all key factors affecting store location selection, including geographic coordinate information, population density within the grid, population age structure, income level, number of commercial facilities, type of commercial facilities, transportation convenience, rent level, distribution of competing stores, urban planning restrictions, traffic flow, and coverage of public transportation and subway stations, etc., totaling twenty core dimensions. Each dimension corresponds to a specific environmental feature indicator, which together constitute a high-dimensional state vector, comprehensively describing the environmental state of the agent. The action space of the agent is clearly defined. The action space refers to all the predefined geospatial grids within the 3D digital twin environment. Each geospatial grid corresponds to a specific candidate store location. Each action of the agent is to select one of these grids as a recommended store location. The action space is kept consistent with the number of geospatial grids to ensure that all potential candidate store locations are included in the agent's decision-making scope. The Xavier initialization method is used to set the initial weights of the reinforcement learning agent's policy network. This method avoids gradient vanishing or exploding problems in the early stages of training by keeping the variance of the input and output of each layer of the network consistent. Specifically, based on the number of input and output neurons in each layer of the policy network, the initial value range of the weights for each layer is calculated. Initial weight values ​​are then randomly generated within this range, completing the initial weight configuration of the policy network. By determining the state space dimension, clarifying the action space, and setting the initial policy network weights, the initial parameters of the reinforcement learning agent are configured, ensuring that the agent possesses preliminary decision-making capabilities and achieving the initial deployment of the agent.

[0075] Step 1.5: Based on the initially deployed reinforcement learning agent and 3D digital twin environment, establish a real-time data connection channel with the urban planning department and traffic management platform, set the data update frequency and change detection threshold, and obtain a dynamic data stream input mechanism. Specifically, this includes: building a dedicated data connection channel between the initially deployed reinforcement learning agent and 3D digital twin environment and the urban planning department and traffic management platform; the connection channel uses an encrypted transmission protocol for data transmission, and through multiple security mechanisms such as identity authentication, data encryption, and transmission verification, ensures the security, integrity, and confidentiality of data during transmission, preventing data theft or... The system prevents data tampering; it sets the data update frequency to once per hour, meaning that the urban planning department and traffic management platform push the latest relevant data to this connection channel every hour to ensure that the intelligent agent obtains environmental change information in a timely manner; it also sets change detection thresholds. For urban planning information, when the change rate of core indicators such as land use, plot ratio, and building density exceeds 5%, it is judged as a significant environmental change, triggering the subsequent environmental update process. For traffic network data, when the change in data such as road traffic flow, bus stop passenger flow, and subway operation frequency exceeds 10%, it is judged as a significant environmental change, triggering the subsequent environmental update process. Through the establishment of the above-mentioned dedicated data connection channel, the setting of the data update frequency, and the clarification of change detection thresholds, a complete dynamic data flow input mechanism is formed.

[0076] Step 1.6 involves continuously receiving updated urban planning and transportation information based on a dynamic data stream input mechanism. This updated data is then integrated with the 3D digital twin environment in real time to generate environmental state change signals. Specifically, this includes: continuously receiving updated data pushed by urban planning departments and transportation management platforms through a dedicated data connection channel, relying on the established dynamic data stream input mechanism; classifying and organizing each received updated data according to urban planning information and transportation network data, then comparing it field-by-field with the corresponding categories of data already stored in the 3D digital twin environment, checking the numerical changes of each data field, and identifying the differences between the updated data and existing data; for the identified discrepancies, integrating them into the 3D digital twin environment in real time according to the data storage specifications and modeling standards of the 3D digital twin environment, synchronously updating the geospatial grid-related attributes, population flow trajectory simulation data, commercial facility distribution association information, and transportation network parameters to ensure that the 3D digital twin environment accurately reflects the current urban situation. After completing data fusion and environmental updates, environmental state change signals are generated based on the state differences of the 3D digital twin environment before and after the update.

[0077] In this embodiment of the invention, by employing multiple channels to acquire urban planning and transportation network data, and processing the data through cleaning, standardization, and spatial coordinate alignment to obtain structured data, a three-dimensional digital twin environment containing elements such as geospatial grids is constructed based on this data. Simultaneously, the initial parameters of the intelligent agent are configured to complete the deployment, and a real-time data connection channel and dynamic data stream input mechanism are established to generate environmental state change signals. This overcomes the technical problems of poor data quality, environmental simulation deviating from the real scene, inability to capture real-time dynamic changes in urban planning and transportation, and insufficient adaptability of intelligent agent deployment to the environment in traditional site selection environment construction. Thus, the accurate construction and dynamic updating of the reinforcement learning basic environment are achieved, ensuring the rationality of intelligent agent deployment.

[0078] In a preferred embodiment of the present invention, step 2 above may include:

[0079] Step 2.1: Based on the environmental status change signals, analyze the impact of urban planning changes and transportation network adjustments to identify the affected urban areas. Specifically, this includes: receiving the generated environmental status change signals and extracting the specific content of urban planning changes and transportation network adjustments, including the scope of land use adjustments involved in the planning changes, the magnitude of plot ratio modifications, the addition or rerouting of transportation routes, station relocation locations, traffic flow adjustment data, and other key information; constructing an assessment index system for the degree of change impact. Primary indicators include the scale of planning changes, the scope of transportation adjustments, the degree of impact on population flow, and the degree of impact on commercial activities. Each primary indicator is further subdivided into secondary indicators. The scale of planning changes covers the percentage of the changed area and the number of core indicator adjustments; the scope of transportation adjustments covers the length of adjusted routes and the number of station changes; the degree of impact on population flow covers the scale of potential passenger flow transfer and the proportion of changes in travel routes; and the degree of impact on commercial activities covers changes in the accessibility of surrounding commercial facilities and changes in the commercial attraction capacity. Experts in urban planning, transportation engineering, and commercial operations were invited to assign weights to each indicator: the scale of planning changes was weighted at 0.3, the scope of transportation adjustments at 0.3, the impact on population flow at 0.2, and the impact on commercial activities at 0.2. Based on the actual data and weights of each indicator, a comprehensive score for the degree of impact of changes was calculated for each urban area. A comprehensive score threshold of 80 points was set, and areas with scores higher than 80 points were identified as the urban areas affected by the changes, ensuring that this range accurately covers the areas affected by urban planning changes and transportation network adjustments.

[0080] Step 2.2: Based on the affected urban area, extract all public transport hub node data from the 3D digital twin environment. Combine this with historical passenger flow data, the number of transfer routes, and the population density coverage index to obtain a candidate public transport hub evaluation list. Specifically, this includes: extracting complete data of all public transport hub nodes within the identified affected urban area from the 3D digital twin environment, including hub name, specific spatial coordinates, construction scale, years of operation, historical passenger flow data, number of transfer routes, population density coverage, and surrounding commercial facilities. The historical passenger flow data is selected from the average daily passenger flow data of the past six months, directly extracted from the traffic operation statistics stored in the 3D digital twin environment; the number of transfer routes is calculated to include the number of transfer routes for various transportation lines such as buses, subways, and long-distance passenger transport that have been opened at the hub, including the number of routes with in-station transfers and convenient out-of-station transfers; the population density coverage is calculated by delineating a circular area with a radius of three kilometers centered on the hub, counting the number of permanent residents and transient residents within this area, and calculating the average population density. The extracted bus hub node data is organized, and the core evaluation indicators are historical passenger flow data, number of transfer routes, and population density. A weighted summation method is used to calculate the comprehensive evaluation score of each bus hub, where historical passenger flow data has a weight of 0.4, the number of transfer routes has a weight of 0.3, and the population density has a weight of 0.3. The bus hubs are sorted from high to low according to the comprehensive evaluation score to form a candidate bus hub evaluation list, ensuring that the bus hubs in the list have a high passenger flow attraction capacity and radiation range.

[0081] Step 2.3: Based on the candidate bus hub evaluation list, a spatial distribution balance algorithm is used to calculate the geometric distance and passenger flow complementarity coefficient between each candidate hub. The three bus hubs with the most balanced spatial distribution and the highest passenger flow coverage complementarity are selected as evaluation base points. Specifically, this includes: retrieving the spatial coordinate data of each candidate hub from the candidate bus hub evaluation list; calculating the geometric distance between each candidate hub using the spatial distribution balance algorithm; obtaining the straight-line distance between any two candidate hubs through coordinate calculation; ensuring that the geometric distance between any two of the three selected hubs is not less than five kilometers to guarantee the spatial distribution balance of the evaluation base points and avoid overlapping coverage or excessively large coverage blind spots; simultaneously... Analyze the historical passenger flow data of each candidate hub, extract passenger flow change curves for different time periods, including morning peak, evening peak, and off-peak periods, and calculate the passenger flow complementarity coefficient between any two candidate hubs. By comparing the peak periods and flow differences of the two passenger flow change curves, the coefficient ranges from zero to one. The closer the coefficient is to one, the greater the difference in peak passenger flow periods between the two hubs, and the stronger the complementarity of passenger flow coverage. Taking into account both the spatial distribution balance and the passenger flow complementarity coefficient, first select candidate hub combinations that meet the geometric distance requirements between each pair, and then select the combination with the highest sum of passenger flow complementarity coefficients from these combinations. The three bus hubs in the combination are the evaluation benchmarks with the most balanced spatial distribution and the highest passenger flow coverage complementarity.

[0082] Step 2.4: Using the spatial coordinates of three public transport hub assessment base points, a triangular coverage area is constructed. By calculating the circumradius of the line connecting the three base points and the regional population density decay function, the boundary of the core business district is determined, resulting in the core business district to be assessed. Specifically, this includes: obtaining the precise spatial coordinates of the three assessment base points; marking the specific locations of the three base points in a 3D digital twin environment based on the coordinate data; constructing a triangular coverage area by connecting the points; calculating the circumradius of this triangle using coordinate calculation tools, i.e., the straight-line distance from the three vertices of the triangle to the center of the circumradius, which serves as the basic coverage reference for the core business district; and constructing the regional population density decay function. ,in, The distance from the center of the circumcircle of the triangle is Population density at the location It represents the initial population density at the center of the circumcircle of the triangle, which is the actual population density at the location of the circumcircle's center. This density is directly obtained from the population distribution data at that location within the 3D digital twin environment. The attenuation coefficient is set to 0.8 to reflect the rate of decrease in population density with increasing distance. It is the straight-line distance from the location where the population density is to be calculated to the center of the circumcircle. The radius of the circumcircle of the triangle, which is the radius of the circumcircle of the triangle formed by connecting the three evaluation base points obtained through coordinate calculation tools, serves as the basic distance reference for population density decline. In the mathematical expression, the function is a power, with a decay coefficient set to 0.8. This coefficient is determined by analyzing the overall pattern of urban population distribution and combining it with the population density characteristics of the affected area. The function shows a trend of gradually decreasing population density from the center of the circumcircle of the triangle towards the surrounding area. Based on the circumcircle of the triangle constructed from three evaluation base points, the function extends gradually from the center to the outer area according to the population density decay function. The population density data at each location is monitored in real time during the extension process. A population density threshold of 500 people per square kilometer is set. When the extension reaches a certain location and the population density at that location is lower than 500 people per square kilometer, that location is determined as the boundary of the core business district. Through the above method, the specific boundary range of the core business district to be evaluated is clarified, ensuring that the area focuses on core areas that are positively affected by dynamic changes, have high population density, and possess good commercial potential.

[0083] In this embodiment of the invention, the impact of urban planning and traffic adjustments on the environmental state change signal is analyzed to identify the affected areas. Data on public transport hubs in these areas is extracted and combined with multiple indicators to form a candidate evaluation list. A spatial distribution balance algorithm is used to select three public transport hubs with spatial balance and high passenger flow complementarity as evaluation base points. Then, a triangular coverage area is constructed based on the base point coordinates, and the core area of ​​the business district is defined by combining the circumcircle radius and the population density decay function. This overcomes the technical problems of traditional business district definition, such as lack of connection to urban dynamic changes, strong subjectivity in the selection of evaluation base points and unbalanced spatial distribution, insufficient passenger flow coverage complementarity, and lack of scientific basis for business district boundary delineation. Therefore, it achieves accurate locking of areas affected by dynamic changes and scientific definition of the core area of ​​the business district, ensuring that the core area of ​​the business district has high commercial potential and evaluation representativeness.

[0084] In a preferred embodiment of the present invention, step 3 above may include:

[0085] Step 3.1: Divide the core area of ​​the business district into regular grids according to a preset resolution to obtain a grid set. Extract the point set data of population distribution, commercial facilities, and transportation stations within each grid from the 3D digital twin environment. Calculate the geometric moments and central moments of the population and commercial facility point sets using a shape moment calculation algorithm to obtain the population distribution concentration parameter and the commercial facility distribution balance parameter. Specifically, this includes: First, dividing the defined core area of ​​the business district into regular grids, setting the preset resolution to 50 meters by 50 meters. Using the boundary coordinates of the core area as a reference, divide the grid row by row and column by column from the boundary starting point to form several square grids of uniform size and clear boundaries. Assign a unique identification code to each grid and clarify the coordinate range of the four corners of each grid to ensure that the grid can fully cover the core area of ​​the business district without overlap or omission, ultimately obtaining a complete grid set. Then, extract three types of point set data from each grid from the 3D digital twin environment: population distribution point set data comes from anonymized mobile phone signaling location data, detailed population census data, and residential community occupancy rate statistics, accurately marking the actual residence location or activity point of each resident within the grid; commercial facility point set data... The data set encompasses the specific location information of all commercial establishments within the grid, including chain stores, shopping malls, supermarkets, restaurants, and entertainment venues, while also linking basic attributes such as the business type and operating area of ​​these commercial facilities. The transportation point set data includes the precise coordinates of various transportation facilities within the grid, such as bus stops, subway entrances / exits, and shared bicycle parking spots, as well as information on operating hours and service areas. A shape moment calculation algorithm is used to calculate the geometric moments of the population distribution point set and commercial facility point set within each grid separately, first calculating the zero-order moment, first-order moment, and second-order moment. The population distribution point set is obtained by summing the total number of points within the set, reflecting its overall size. The first moment is obtained by averaging the coordinates of all points, reflecting the central location of the set. The second moment is obtained by calculating the dispersion of all points relative to the origin, reflecting the spread of the set. The central moments, including the first and second moments, are then calculated. The first moment is obtained by summing the offsets of all points relative to the center of the set, and the second moment is obtained by summing the squares of the offsets of all points relative to the center, reflecting the concentration of the points around the center. The population distribution concentration parameter is obtained by the ratio of the second to the zeroth moment of the population distribution point set; a larger ratio indicates a more concentrated population distribution. The distribution range of the second central moment of the commercial facility point set provides a parameter for the balance of commercial facility distribution; a smaller range indicates a more balanced distribution of commercial facilities.

[0086] Step 3.2: Calculate the moment characteristics of the traffic station distribution point set in each grid using the population distribution concentration parameter and the commercial facility distribution balance parameter. This includes the compactness parameter and coverage uniformity parameter of the station distribution. Specifically, based on the obtained population distribution concentration parameter and commercial facility distribution balance parameter of each grid, calculate the moment characteristics of the traffic station distribution point set in each grid. First, the calculation range of the traffic station distribution point set is ensured to be completely consistent with the grid boundary to guarantee the relevance of the calculation results. When calculating the compactness parameter of the station distribution, the minimum bounding rectangle of the traffic station distribution point set is obtained first through a shape moment calculation algorithm. This rectangle can just accommodate all traffic stations and its edges are parallel to the coordinate axes. The area of ​​this minimum bounding rectangle is calculated, and then the actual area covered by the traffic stations is calculated, i.e., the total area covered by the 50-meter service radius around all stations. The compactness parameter is obtained by the ratio of the area of ​​the minimum bounding rectangle to the actual coverage area. The closer the ratio is to one, the more compact the traffic station distribution and the higher the overlap of the service range. When calculating the coverage uniformity parameter of the station distribution, the first-order central moment of the traffic station distribution point set is obtained through a shape moment calculation algorithm to determine the center position of the point set. Then, the straight-line distance from each traffic station to this center position is calculated, and the variance of all distances is calculated. The smaller the variance value, the more uniform the distribution of traffic stations around the center position and the more balanced the coverage range. Through the above calculations, the corresponding compactness parameter and coverage uniformity parameter are obtained for the traffic station distribution within each grid.

[0087] Step 3.3: Based on the concentration parameter, balance parameter, and traffic station moment characteristics, the population density index, commercial activity index, and transportation convenience index of each grid are comprehensively calculated using a multi-dimensional fusion evaluation method to obtain the grid comprehensive evaluation vector. Specifically, the multi-dimensional fusion evaluation method is adopted to first clarify the weight allocation rules of the population density index, commercial activity index, and transportation convenience index. The weight of the population density index is set to 0.4, the weight of the commercial activity index is set to 0.3, and the weight of the transportation convenience index is set to 0.3, ensuring that the sum of the weights of the three indicators is one, thus guaranteeing the balance and rationality of the evaluation dimensions. When calculating the population density index, the initial population density is first obtained by dividing the zero-order geometric moment of the population distribution point set within the grid by the grid area. Then, it is corrected by combining the population distribution concentration parameter. The concentration parameter judgment threshold is set to 0.8 and 1.2. If the concentration parameter is greater than 1.2, it indicates that the population distribution within the grid is highly concentrated, and the initial population density is multiplied by a correction factor of 1.1. If the concentration parameter is less than 0.8, it indicates that the population distribution within the grid is relatively dispersed, and the initial population density is multiplied by a correction factor of 0.9. If the concentration parameter is between 0.8 and 1.2, it indicates that the population distribution concentration is moderate, and the initial population density remains unchanged. When calculating the commercial activity index, the initial commercial density is first obtained by dividing the zero-order geometric moment of the set of commercial facility points within the grid by the grid area. Then, a correction is made based on the commercial facility distribution balance parameter. The balance parameter thresholds are set at 0.6 and 1.4. If the balance parameter is less than 0.6, it indicates a highly balanced distribution of commercial facilities, and the initial commercial density is multiplied by a correction factor of 1.05. If the balance parameter is greater than 1.4, it indicates an uneven distribution of commercial facilities, and the initial commercial density is multiplied by a correction factor of 0.95. If the balance parameter is between 0.6 and 1.4, it indicates a moderate degree of balance in the distribution of commercial facilities, and the initial commercial density remains unchanged. When calculating the transportation convenience index, the initial density is first obtained by dividing the set of commercial facility points within the grid by the grid area. The initial station density is obtained by dividing the zero-order geometric moment of the distribution set of internal transportation stations by the grid area. This density is then jointly corrected using the station compactness and coverage uniformity parameters. The compactness parameter thresholds are set to 0.65 and 0.85, and the coverage uniformity parameter thresholds are 200 and 400. If the compactness parameter is greater than 0.85 and the coverage uniformity parameter is less than 200, the transportation stations are considered compactly distributed and evenly covered, and the initial station density is multiplied by a correction factor of 1.1. If the compactness parameter is less than 0.65 or the coverage uniformity parameter is greater than 400, the transportation stations are considered loosely distributed or unevenly covered, and the initial station density is multiplied by a correction factor of 0.9. Otherwise, the initial station density remains unchanged. After calculating and correcting the three indicators, the final score of each indicator is multiplied by its corresponding weight, and the three weighted scores are summed to obtain the comprehensive evaluation score for each grid. The population density score, commercial activity score, transportation convenience score, and comprehensive evaluation score for each grid together constitute a four-dimensional grid comprehensive evaluation vector.

[0088] Step 3.4: Based on the grid comprehensive evaluation vector, the regional characteristic compensation parameter matrix for calibrating the original monitoring data is obtained through regional characteristic difference analysis and standardization. Specifically, this includes: performing regional characteristic difference analysis on the comprehensive evaluation vectors of all grids; first, calculating the overall average values ​​of all grids within the core business district for population density, commercial popularity, and transportation convenience indicators; then, calculating the difference between the scores of each grid for the three indicators and the corresponding overall average value. A positive difference indicates that the grid performs better than the overall level for that indicator, while a negative difference indicates that the grid performs worse than the overall level for that indicator; and performing standardization based on the difference results, setting the range of the standardized compensation parameters to 0.7 to... 1.3. The index difference of each grid is proportionally converted to this range. If the population density index difference of a certain grid is positive and has the largest absolute value, the compensation parameter corresponding to that index is set to 1.3. If the difference is negative and has the largest absolute value, the compensation parameter is set to 0.7. The compensation parameters of other grids are calculated linearly according to the size of the difference. Each grid corresponds to three compensation parameters, which correspond to the three indices of population density, commercial popularity, and transportation convenience. The compensation parameters of all grids are arranged in the order of grid identification codes to form a two-dimensional regional feature compensation parameter matrix. The number of rows in the matrix is ​​the same as the number of grids, and the number of columns is three. Each element corresponds to the compensation parameter of a specific index of a specific grid.

[0089] In this embodiment of the invention, by employing a regular grid division of the core area of ​​the business district, extracting point set data of population, commercial facilities, and transportation stations within each grid, quantifying the distribution characteristic parameters of various point sets through a shape moment calculation algorithm, and combining multiple parameters through a multi-dimensional fusion evaluation model to obtain a grid comprehensive evaluation vector, and then generating a regional feature compensation parameter matrix through regional feature difference analysis and standardization processing, the technical problems of coarse division, inaccurate quantification of population, commercial and transportation distribution characteristics, imbalance of multi-dimensional indicator fusion, and lack of targeted data calibration basis in traditional regional feature evaluation are overcome. Thus, a refined and accurate quantitative evaluation of the spatial characteristics of the core area of ​​the business district is achieved, and a reliable regional feature compensation parameter matrix is ​​obtained.

[0090] In a preferred embodiment of the present invention, step 4 above may include:

[0091] Step 4.1 involves performing grid-by-grid calibration on the structured urban spatial base data based on the regional feature compensation parameter matrix to obtain calibrated spatial data. Specifically, this includes retrieving the regional feature compensation parameter matrix, where each element corresponds to a specific index compensation parameter for a particular grid. First, a one-to-one correspondence is established between the compensation parameter matrix and the grid set of the core business district through grid identification coding. Each grid can be matched with corresponding population density compensation parameters, commercial popularity compensation parameters, and transportation convenience compensation parameters. Then, grid-by-grid calibration is carried out. For each grid, the corresponding three compensation parameters are multiplied by the population distribution statistics, the number and scale statistics of commercial facilities, and the service capacity statistics of transportation stations in the structured urban spatial base data for that grid, respectively, to obtain the calibrated data for each field. For example, if the original population density statistics within a grid are 6,000 people per square kilometer, and the corresponding population density compensation parameter is 1.1, then the calibrated population density is 6,600 people per square kilometer. If the commercial heat compensation parameter of a certain grid is 0.9, and the original commercial facility revenue scale statistics are 2 million yuan per year, then the calibrated revenue scale statistics are 1.8 million yuan per year. By using the above method, the accurate calibration of all grid core data is completed, and the calibrated spatial data is obtained.

[0092] Step 4.2 involves data fusion and smoothing of the calibrated spatial data to obtain the corrected monitoring results. Specifically, this includes: multi-source data fusion processing of the calibrated spatial data, linking and integrating the calibrated population, commercial, and transportation data within the same grid by field to form a complete data archive for each grid, ensuring data relevance and integrity; subsequently, data smoothing processing is performed, using a moving average method to eliminate data abrupt changes at grid boundaries, with the moving window size set to three adjacent grids. This means the final data for each grid is calculated from its own data and the calibration data of the two adjacent grids (left, right, top, and bottom). The current grid data weight is set to 0.5, and the weights of the two adjacent grid data are each set to 0.25. The smoothed single-grid data is obtained through weighted summation. Fusion and smoothing processing are then performed on all grids sequentially to finally obtain the corrected monitoring results.

[0093] Step 4.3 involves dynamically updating the constructed 3D digital twin environment based on the corrected monitoring results to obtain updated environmental data. Specifically, this includes initiating the dynamic update process for the 3D digital twin environment using the corrected monitoring results as the data source. First, the basic spatial data layer of the environment is updated, replacing all existing data such as population distribution points, commercial facility models, and transportation station layouts with the latest corrected data to ensure that the location and attributes of spatial entities are consistent with the actual situation. Next, the dynamic simulation layer is updated. Based on the corrected population data, the simulation parameters for population flow trajectories are adjusted to better reflect real-world passenger flow changes. The simulation of the operational status of commercial facilities, including operating hours and passenger attraction capacity, is updated based on the corrected commercial data. The simulation of traffic flow is adjusted based on the corrected traffic data to ensure that simulation results such as route congestion and station passenger flow accurately reflect the actual traffic conditions. During the update process, the accuracy of data fusion is verified in real time to avoid data conflicts or omissions. After updating all data layers, the updated environmental data is obtained.

[0094] Step 4.4: Extract key state features from the updated environmental data to obtain the current environmental state. Specifically, this includes: extracting key state features from the updated environmental data that comprehensively reflect the state of the core business district. The feature dimensions cover eighteen core elements, including the calibrated population density of each grid, the proportion of population age structure, the number and proportion of commercial facilities and business types, the commercial popularity index, the number and coverage of transportation stations, peak and off-peak traffic data, rent levels, the number and scale of competing stores in the surrounding area, urban planning restrictions, and regional economic activity. All extracted feature data are standardized, and all feature values ​​are uniformly converted to the range of zero to one to eliminate the difference in the dimensions of different data dimensions. For example, the population density range from hundreds to tens of thousands per square kilometer is proportionally mapped to a standardized value of zero to one. Finally, all standardized feature data are arranged in a preset order to form a high-dimensional current environmental state vector.

[0095] In this embodiment of the invention, the structured urban spatial basic data is precisely calibrated grid by grid based on a regional feature compensation parameter matrix. The calibrated data is then fused and smoothed to eliminate data deviations. The three-dimensional digital twin environment is dynamically updated based on the corrected monitoring results. Key state features are extracted from the updated environmental data to obtain the current environmental state. This overcomes the technical problems of traditional artificial intelligence, such as coarse calibration methods, data mutation interference, disconnect between environmental updates and actual states, and inaccurate extraction of key state features. As a result, the accurate correction of urban spatial data and the dynamic synchronous update of the reinforcement learning environment are achieved, ensuring that the current environmental state can truly and comprehensively reflect the actual situation of the core business district.

[0096] In a preferred embodiment of the present invention, step 5 above may include:

[0097] Step 5.1: By receiving the current environmental state, input the current environmental state into the initialized reinforcement learning agent; based on the current environmental state, call the adaptive weighted reward function built into the reinforcement learning agent, and dynamically adjust the weight ratio of the three indicators of rental cost, expected customer flow, and return on investment in the reward calculation according to the changing trend of urban economic activity and historical operating data. Specifically, this includes: first, receiving the current environmental state, which is presented in the form of a high-dimensional vector, containing standardized data of 18 core elements such as calibrated population density, commercial popularity, transportation convenience, rental level, distribution of competing stores, and urban economic activity in each grid of the core business district; inputting this high-dimensional state vector completely into the initialized reinforcement learning agent to ensure that the agent fully obtains the real-time environmental information of the core business district; then calling the adaptive weighted reward function built into the agent. ,in, For the overall reward value, The rental cost indicator is dynamically weighted. The dynamic weighting of expected passenger flow indicators The weighting of the rate of return on investment is dynamically adjusted according to the trend of urban economic activity. Standardize the score for the rental cost indicator. Standardized score for expected passenger flow indicator, To standardize the return on investment (ROI) score, a dynamic adjustment process for indicator weights was initiated. First, data on urban economic activity was obtained from local statistical department databases and business association reports, including the GDP growth rate, total retail sales growth rate, and per capita disposable income growth rate for the past three quarters. The average of these three indicators was calculated to obtain the comprehensive economic activity growth rate. Thresholds for the growth rate were set at 5% and 2%. If the comprehensive growth rate was greater than 5%, the urban economic activity was considered to be on an upward trend; if it was less than 2%, it was considered to be on a downward trend; and if it was between the two, it was considered to be on a stable trend. Simultaneously, the brand's internal historical operational database was retrieved to extract data on average daily customer traffic, rental expenses, actual ROI, and operating efficiency for similar stores under different environmental conditions over the past two years. The study analyzes the correlation between three indicators—rental cost, expected customer traffic, and return on investment (ROI)—and final operating efficiency. Based on the analysis results, the weighting ratios are adjusted, with the initial weighting ratios set at 30% for rental cost, 40% for expected customer traffic, and 30% for ROI. If economic activity shows an upward trend, and historical data shows a higher correlation between expected customer traffic and ROI and operating efficiency, the weighting of expected customer traffic is increased to 45%, ROI to 35%, and rental cost to 20%. If economic activity shows a downward trend, and historical data shows a more significant impact of rental cost control on operating efficiency, the weighting of rental cost is increased to 40%, expected customer traffic to 35%, and ROI to 25%. Under stable trends, the initial weightings are maintained, or the weighting ratios are slightly adjusted by no more than 5 percentage points based on the correlation of historical data.

[0098] Step 5.2: Based on dynamically adjusted weight ratios and the current environmental status, conduct multi-dimensional evaluations of all candidate geographic grid locations within the core business district, and calculate the comprehensive reward score for each candidate location. Specifically, this includes: based on dynamically adjusted weight ratios and the current environmental status, all regularized grids within the core business district are considered candidate geographic grid locations, and multi-dimensional evaluations are conducted for each. For each candidate grid, corresponding dimension data is extracted from the current environmental status vector to calculate scores for three core indicators: the expected passenger flow indicator score is determined based on data such as calibrated population density within the grid, uniformity of transportation station coverage, and the number of surrounding residential communities. Higher population density, more uniform transportation coverage, and more residential communities result in higher scores. The data is standardized and converted into a score range of 0 to 100. For example, a population density of 8000 people per square kilometer and a transportation coverage uniformity parameter less than 200 receive a full score of 100. The rental cost indicator score is calculated based on rental level data within the grid. The lower the rent level, the higher the score, also standardized to 0 to 100 points. For example, a rent below 80 yuan per square meter per month gets 100 points, while a rent above 200 yuan per square meter per month gets less than 30 points. The return on investment (ROI) score is calculated by combining the expected revenue, rental costs, and other operating costs corresponding to the expected customer flow. Expected revenue is obtained by multiplying the expected customer flow by the average spending per person. Operating costs include staff salaries, utilities, and raw material procurement costs. ROI equals expected revenue minus total costs divided by total investment, standardized to 0 to 100 points based on the ROI value. For example, a ROI above 15% gets 100 points, while a ROI below 5% gets less than 30 points. After calculating the scores for the three indicators, each indicator score is multiplied by its corresponding dynamic weight ratio to obtain a weighted score for each indicator. The three weighted scores are then added together to obtain the comprehensive reward score for each candidate geographic grid location, ranging from 0 to 100 points. A higher score indicates higher feasibility and potential operating benefits for the location.

[0099] Step 5.3: Based on the comprehensive reward score, candidate locations are sorted and screened. Combined with regional competition analysis and brand positioning requirements, the final site selection strategy is derived, including recommended store location coordinates and expected operating efficiency indicators. Specifically, this includes: First, all candidate geographic grid locations are sorted from highest to lowest based on their comprehensive reward score, setting a score screening threshold of 80 points. Candidate locations with scores higher than 80 points are selected to form a preliminary candidate list, ensuring that the locations on the list have high basic commercial potential. Then, a regional competition analysis is conducted. For each location in the preliminary candidate list, data such as the number of similar chain stores and local similar stores within a 3-dimensional digital twin environment, operating scale, years of operation, average daily customer traffic, and customer reviews are extracted from the 3D digital twin environment and commercial supervision database. A regional competition intensity index is calculated, setting competition intensity thresholds of 3.5 and 6.5 points, with a maximum score of 10. Areas with a competition intensity index below 3.5 points are considered low-competition areas and are retained; areas above 6.5 points are considered high-competition areas and are directly eliminated. For areas between these two thresholds, further evaluation is conducted based on the store's differentiation advantages. If the brand has significant differences in products and services... If a location is deemed suitable, it is retained; otherwise, it is eliminated. A second screening process is then conducted based on brand positioning requirements. If the brand is positioned as a community convenience store, priority is given to candidate locations with densely populated residential areas and a population density, requiring an occupancy rate of over 80% in the surrounding residential communities. If positioned as a commercial core store, priority is given to candidate locations with high commercial activity and high traffic volume, requiring at least 15 commercial facilities within the grid and an average daily traffic flow of over 5,000 people. If positioned as a high-end boutique store, priority is given to candidate locations with a high concentration of high-income individuals and well-developed high-end commercial facilities nearby, requiring a per capita disposable income within the grid to be 50% higher than the city average. After competitive landscape analysis and brand positioning matching, the remaining candidate locations with the highest comprehensive reward score are selected as the final store location. The precise coordinates are determined, i.e., the coordinate range of the four corners of the grid. Simultaneously, based on the current environmental conditions and historical operating data, expected operating efficiency indicators are calculated, including expected daily customer traffic, expected monthly revenue, expected monthly operating costs, expected return on investment, and expected investment recovery period. The final store location coordinates and the aforementioned expected operating efficiency indicators are integrated to form the final site selection strategy.

[0100] In this embodiment of the invention, the technical means of receiving the current environmental state and inputting it into a reinforcement learning agent, calling an adaptive weighted reward function to dynamically adjust the weight ratios of rental costs, expected customer flow, and return on investment based on the changing trends of urban economic activity and historical operating data, and then performing multi-dimensional evaluation of candidate geographic grid locations based on the adjusted weights and environmental state to calculate a comprehensive reward score, and then combining regional competitive situation analysis and brand positioning requirements to obtain a final site selection strategy that includes recommended store locations and expected operating efficiency indicators, overcomes the technical problems of traditional site selection decisions, such as fixed multi-objective weight settings, inability to adapt to dynamic market changes, single evaluation dimensions of candidate store locations, and insufficient consideration of the competitive environment and brand positioning leading to one-sided decisions. This achieves dynamic adaptability and multi-dimensional comprehensiveness in site selection evaluation, ensuring that the final site selection strategy is both in line with actual market needs and brand development positioning, thus improving the accuracy and rationality of store location selection.

[0101] In a preferred embodiment of the present invention, step 6 above may include:

[0102] Step 6.1 involves receiving the site selection strategy and sending new store opening instructions and recommended store location coordinates to store management. This includes: first, receiving the final site selection strategy, which contains the precise coordinates of the final store location (the coordinate range of the four corners of the corresponding grid), as well as expected operating efficiency indicators such as expected daily customer traffic, expected monthly revenue, expected monthly operating costs, expected return on investment, and expected investment payback period. Based on the received site selection strategy, a standardized new store opening instruction is generated. This instruction clearly indicates the recommended store location's coordinate range, the store preparation timeline, the planned store area, and basic decoration standards. Through the brand's internal store management information platform, the new store opening instruction and recommended store location coordinates are simultaneously sent to the store expansion department, engineering preparation department, and operations management department. A confirmation mechanism for instruction delivery is also established, requiring each receiving department to provide feedback on the receipt status and preliminary execution plan within 24 hours to ensure that the opening instruction is accurately and efficiently transmitted to each execution stage.

[0103] Step 6.2: Based on the execution results of the new store opening instruction, collect real-time operational return data of the new store, including average daily customer traffic, single-store revenue, and cost expenditure data. Specifically, after the new store opening instruction is completed and the store officially opens, collect the actual operational return data of the new store hourly. The average daily customer traffic data is obtained by averaging the data collected from infrared sensing devices at the store entrance and video surveillance counting devices to ensure data accuracy. The single-store revenue data is directly extracted from the POS terminal, including detailed data such as total daily revenue, revenue by time period, and revenue of various product categories. The cost expenditure data is obtained from the financial accounting module and energy consumption metering device, and is broken down into specific items such as rent, personnel salary, water and electricity, raw material procurement costs, and equipment maintenance costs. All collected operational return data is summarized and organized daily and stored in the brand's dedicated operational data center. At the same time, a data anomaly warning mechanism is set up. When the daily fluctuation of a certain data item exceeds 20%, an alert is automatically triggered and the operational management personnel are notified to verify, ensuring that the collected data is true, complete, and continuous.

[0104] Step 6.3: Based on actual operational return data, calculate operational performance indicators (RPIs) and use them as environmental feedback signals. Specifically, this includes: calculating RPIs based on collected actual operational return data; the key performance indicators to be calculated include four key metrics: actual return on investment (ROI), customer conversion rate, cost control rate, and revenue achievement rate. ROI is calculated by dividing the difference between actual monthly revenue and actual total monthly cost by the total investment amount; customer conversion rate is calculated by dividing the actual average daily number of customers by the actual average daily customer flow; cost control rate is calculated by dividing the actual total monthly cost by the given expected monthly operating cost; and revenue achievement rate is calculated by dividing the actual monthly revenue by the expected monthly revenue. The four core performance indicators are standardized by converting all indicator values ​​into a scoring system from 0 to 100 points. For example, an ROI of 15% or higher receives 100 points, while a ROI below 5% receives less than 30 points; a customer conversion rate above 30% receives 100 points, while a customer conversion rate below 10% receives less than 30 points. The standardized scores of the four performance indicators are integrated into an environmental feedback signal.

[0105] Step 6.4: Based on the environmental feedback signals, associate the corresponding environmental states and site selection strategies to construct a complete training sample dataset. This includes: first, extracting standardized actual operational performance index scores from the environmental feedback signals; then, retrieving the current environmental state (i.e., the generated high-dimensional environmental state vector) corresponding to the new store's site selection process, and the corresponding final site selection strategy (i.e., the output store location coordinates and expected benefit indicators); establishing data association rules according to the correspondence between environmental states, site selection strategies, and environmental feedback; and accurately matching the three through timestamps and unique store location identifiers to ensure a one-to-one correspondence between the environmental input and operational feedback for each site selection decision. The three sets of matched data are used as a complete training sample, and the samples are summarized and organized weekly to construct the training sample dataset. Simultaneously, the samples in the dataset undergo quality screening, removing initial operational data from the first 30 days before the new store's opening and abnormal operational data during major holidays and public emergencies to ensure the validity and representativeness of the samples within the dataset, ultimately forming a training sample dataset that truly reflects the correlation between site selection decisions and actual operational results.

[0106] Step 6.5 involves training the initialized reinforcement learning agent using the training sample dataset at a preset time interval. Specifically, this includes: setting the preset training period to 3 months, meaning that the initialized reinforcement learning agent will undergo parameter update training every 3 months using the training sample dataset; before training, dividing the training sample dataset into a training set and a validation set in a 7:3 ratio, with the training set used for parameter adjustment and the validation set used for evaluating training effectiveness; during training, inputting environmental state data from the training set into the reinforcement learning agent, allowing the agent to re-output the corresponding location decision, and then comparing this decision with the actual location strategy and environmental feedback in the dataset to calculate the decision bias; adjusting the weight parameters of the agent's policy network based on the decision bias value, using gradient descent to gradually optimize the parameters until the decision bias value on the validation set is below 5%, at which point training stops. The parameter adjustment trajectory and bias value change curve are recorded in real time during training, and a training log is established to ensure the training process is traceable and reviewable, while avoiding overfitting.

[0107] Step 6.6: Based on the parameter-updated training of the reinforcement learning agent, obtain the optimized decision-making strategy. Specifically, this includes: after completing the parameter update training of the reinforcement learning agent, performing performance testing on the trained agent; selecting site selection case data of new stores in other cities within the past 6 months that did not participate in the training as a test set, inputting the environmental state of the test set into the trained agent, allowing it to output a site selection strategy, and then comparing the strategy with the actual site selection results and operational effects; if the test results show that the deviation between the expected operating efficiency index and the actual operating effect corresponding to the site selection strategy output by the agent is reduced by more than 15% compared with before training, then the training is deemed effective; based on the parameter-updated training of the reinforcement learning agent that has passed the test, generate the optimized decision-making strategy.

[0108] In this embodiment of the invention, by adopting a method of receiving site selection strategies and sending opening instructions and recommended store location coordinates to store management, collecting real-time actual operational return data of new stores through a store operation monitoring system, calculating operational performance indicators based on this data as environmental feedback signals, constructing a complete training sample dataset by associating the corresponding environmental state with the site selection strategy, and updating and training the reinforcement learning agent's parameters according to a preset time period to obtain an optimized decision-making strategy, this invention overcomes the technical problems of traditional site selection methods, such as lack of actual operational data feedback mechanisms, inability to iteratively optimize decision-making strategies based on real operating results, and difficulty in continuously adapting to market changes. This forms a closed-loop mechanism of decision-making, execution, feedback, and optimization, and achieves continuous improvement in the decision-making ability of the reinforcement learning agent.

[0109] like Figure 2 As shown, embodiments of the present invention also provide a chain store site selection optimization system based on reinforcement learning, including:

[0110] The monitoring module is used to build a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time.

[0111] The delineation module is used to select three public transport hubs as evaluation base points in the urban geospatial space based on the monitored change signals, and to jointly delineate the core area of ​​the business district to be evaluated using the three public transport hub base points.

[0112] The partitioning module is used to divide the core area of ​​the business district into grids. Based on the geometric moments and central moments of the population distribution point set and the commercial facility point set within each grid unit, it quantifies the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, it comprehensively evaluates the population density, commercial popularity and transportation convenience indicators within each grid to obtain regional characteristic compensation parameters.

[0113] The correction module is used to calibrate the original monitoring data using regional feature compensation parameters to obtain corrected monitoring results; and to update the reinforcement learning environment based on the monitoring results to obtain the current environmental state.

[0114] The calculation module is used to input the current environmental state into the reinforcement learning agent; the agent dynamically adjusts the weights of rental cost, expected customer flow and return on investment in the reward calculation by calling the adaptive weighted reward function, and evaluates the candidate store sites in combination with the environmental state to obtain the site selection strategy.

[0115] The processing module is used to open new stores through site selection strategies and collect actual operational return data. The operational return data is used as environmental feedback, which, together with the corresponding environmental state and site selection strategy, constitutes training samples to periodically train the reinforcement learning agent in order to optimize the decision-making strategy.

[0116] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A chain store site selection optimization method based on reinforcement learning, characterized in that, The method includes: Construct a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time; Based on the monitored change signals, three public transport hubs were selected in the urban geospatial space as evaluation base points, and the core area of ​​the business district to be evaluated was defined by the three public transport hub base points. By dividing the core area of ​​the business district into grids, the geometric moments and central moments of the population distribution point set and commercial facility point set within each grid unit are used to quantify the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, the population density, commercial popularity and transportation convenience indicators within each grid are comprehensively evaluated to obtain regional characteristic compensation parameters. The original monitoring data is calibrated using regional feature compensation parameters to obtain corrected monitoring results; the reinforcement learning environment is updated based on the monitoring results to obtain the current environmental state. The current environmental state is input into the reinforcement learning agent; the agent dynamically adjusts the weights of rental costs, expected customer flow, and return on investment in the reward calculation by calling an adaptive weighted reward function, and evaluates candidate geographic grid locations in conjunction with the current environmental state to obtain a site selection strategy, including: By receiving the current environmental state, the current environmental state is input into the initialized reinforcement learning agent; based on the current environmental state, the adaptive weighted reward function built into the reinforcement learning agent is called, and the weight ratio of the three indicators of rental cost, expected passenger flow and investment return rate in the reward calculation is dynamically adjusted according to the changing trend of urban economic activity and historical operating data. Based on the dynamically adjusted weight ratio and the current environmental state, all regularized grids in the core area of ​​the business district are taken as candidate geographic grid locations. For each candidate grid, the corresponding dimension data is extracted from the current environmental state vector to calculate the scores of three core indicators. After the scores of the three indicators are calculated, each indicator score is multiplied by the corresponding dynamic weight ratio to obtain the weighted score of each indicator. Then, the three weighted scores are added together to obtain the comprehensive reward score of each candidate geographic grid location. Candidate geographic grid locations are sorted and filtered based on comprehensive reward scores. Combined with regional competitive landscape analysis and brand positioning requirements, the final site selection strategy is obtained, including recommended store location coordinates and expected operating efficiency indicators. By opening new stores through site selection strategies, actual operational return data is collected. This operational return data is used as environmental feedback, which, together with the corresponding environmental conditions and site selection strategies, constitutes training samples for periodic training of the reinforcement learning agent to optimize decision-making strategies.

2. The chain store site selection optimization method based on reinforcement learning according to claim 1, characterized in that, Construct a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time, including: Urban planning information and traffic network data are obtained from urban planning databases, data interfaces of traffic management departments, and third-party geographic information to form the original monitoring dataset; By cleaning, standardizing the format, and aligning the spatial coordinates of the original monitoring dataset, structured urban spatial basic data is obtained. A three-dimensional digital twin environment, including geospatial grids, population flow trajectories, and the distribution of commercial facilities, is constructed based on structured urban spatial data as a foundational environmental framework for reinforcement learning. The initial parameters of the reinforcement learning agent are configured based on the three-dimensional digital twin environment, including the state space dimension, action space range and initial policy network weights, to complete the initial deployment of the agent; Based on the reinforcement learning agent and 3D digital twin environment deployed in the initial stage, a real-time data connection channel is established with the urban planning department and traffic management platform. The data update frequency and change detection threshold are set to obtain a dynamic data flow input mechanism. Based on a dynamic data stream input mechanism, the system continuously receives updated urban planning and traffic information, integrates the updated data with the three-dimensional digital twin environment in real time, and generates environmental status change signals.

3. The chain store site selection optimization method based on reinforcement learning according to claim 2, characterized in that, Based on the monitored change signals, three public transport hubs were selected as evaluation benchmarks in the urban geographic space. These three public transport hubs jointly define the core area of ​​the business district to be evaluated, including: Based on signals of changes in environmental conditions, we analyze the degree of impact of changes in urban planning and transportation network adjustments, and identify the urban areas affected by these changes. Based on the urban area affected by the changes, data on all public transport hub nodes within the area are extracted from the three-dimensional digital twin environment. Combined with historical passenger flow data, number of transfer lines, and population density indicators of the hubs, a candidate public transport hub evaluation list is obtained. Based on the candidate bus hub evaluation list, the spatial distribution balance algorithm is used to calculate the geometric distance and passenger flow complementarity coefficient between each candidate hub, and the three bus hubs with the most balanced spatial distribution and the highest passenger flow coverage complementarity are selected as evaluation base points. By assessing the spatial coordinates of three public transport hubs as evaluation base points, a triangular coverage area is constructed. By calculating the radius of the circumcircle connecting the three base points and the regional population density decay function, the boundary range of the core business district is determined, thus obtaining the core business district to be evaluated.

4. The chain store site selection optimization method based on reinforcement learning according to claim 3, characterized in that, By dividing the core area of ​​the business district into grids, and based on the geometric moments and central moments of the population distribution point set and commercial facility point set within each grid unit, the distribution concentration and dispersion balance characteristics of various point sets are quantified. Simultaneously, combined with the moment characteristics of the transportation station distribution point set within the grid, the population density, commercial activity, and transportation convenience indicators within each grid are comprehensively evaluated to obtain regional characteristic compensation parameters, including: The core area of ​​the business district is divided into regular grids according to a preset resolution to obtain a grid set. The point set data of population distribution, commercial facilities and transportation stations in each grid are extracted from the three-dimensional digital twin environment. The geometric moments and central moments of the population and commercial facility point sets are calculated by the shape moment calculation algorithm to obtain the population distribution concentration parameter and the commercial facility distribution balance parameter. The moment characteristics of the traffic station distribution point set within each grid are calculated using the population distribution concentration parameter and the commercial facility distribution balance parameter, including the station distribution compactness parameter and coverage uniformity parameter. Based on concentration parameters, balance parameters, and traffic station moment characteristics, a multi-dimensional fusion evaluation method is used to comprehensively calculate the population density index, commercial popularity index, and transportation convenience index of each grid, resulting in a grid comprehensive evaluation vector. Based on the grid-based comprehensive evaluation vector, the regional feature compensation parameter matrix for calibrating the original monitoring data is obtained through regional feature difference analysis and standardization processing.

5. The chain store site selection optimization method based on reinforcement learning according to claim 4, characterized in that, The original monitoring data was calibrated using regional feature compensation parameters to obtain corrected monitoring results; The reinforcement learning environment is updated based on the monitoring results to obtain the current environment state, including: The structured urban spatial baseline data is calibrated grid by grid based on the regional feature compensation parameter matrix to obtain calibrated spatial data; By performing data fusion and smoothing on the calibrated spatial data, the corrected monitoring results are obtained; The constructed three-dimensional digital twin environment is dynamically updated based on the corrected monitoring results to obtain updated environmental data; The current environmental state is obtained by extracting key state features from the updated environmental data.

6. The chain store site selection optimization method based on reinforcement learning according to claim 5, characterized in that, New stores are opened through site selection strategies, and actual operational return data is collected. This operational return data, along with the corresponding environmental conditions and site selection strategies, forms training samples for periodic training of the reinforcement learning agent to optimize decision-making strategies, including: By receiving site selection strategies, instructions for opening new stores and recommended store location coordinates are sent to store management. Based on the execution results of the new store opening instructions, the actual operational return data of the new stores is collected in real time, including average daily customer traffic, single store revenue and cost expenditure data; Based on actual operational return data, operational performance indicators are calculated and used as environmental feedback signals. Based on environmental feedback signals, the corresponding environmental states and location strategies are associated to construct a complete training sample dataset; Using the training sample dataset, the parameters of the initialized reinforcement learning agent are updated and trained according to a preset time period. The optimized decision-making strategy is obtained by updating the parameters of the trained reinforcement learning agent.

7. A chain store site selection optimization system based on reinforcement learning, wherein the system implements the method as described in any one of claims 1 to 6, characterized in that, include: The monitoring module is used to build a reinforcement learning environment for store site selection and initialize a reinforcement learning agent to monitor changes in external urban planning and traffic information in real time. The delineation module is used to select three public transport hubs as evaluation base points in the urban geospatial space based on the monitored change signals, and to jointly delineate the core area of ​​the business district to be evaluated using the three public transport hub base points. The partitioning module is used to divide the core area of ​​the business district into grids. Based on the geometric moments and central moments of the population distribution point set and the commercial facility point set within each grid unit, it quantifies the distribution concentration and dispersion balance characteristics of various point sets. At the same time, combined with the moment characteristics of the transportation station distribution point set within the grid, it comprehensively evaluates the population density, commercial popularity and transportation convenience indicators within each grid to obtain regional characteristic compensation parameters. The correction module is used to calibrate the original monitoring data using regional feature compensation parameters to obtain corrected monitoring results; and to update the reinforcement learning environment based on the monitoring results to obtain the current environmental state. The computation module is used to input the current environmental state into the reinforcement learning agent; The agent dynamically adjusts the weights of rental costs, expected customer flow, and return on investment in the reward calculation by calling an adaptive weighted reward function, and evaluates candidate geographic grid locations in combination with the current environmental status to obtain a site selection strategy; The processing module is used to open new stores through site selection strategies and collect actual operational return data. The operational return data is used as environmental feedback, which, together with the corresponding environmental state and site selection strategy, constitutes training samples to periodically train the reinforcement learning agent in order to optimize the decision-making strategy.

8. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • CN117726369A

  • CN119444309A