Dynamic site selection method and device based on enhanced deep learning, and program product

By using a reinforcement deep learning-based method and multi-dimensional data and a site selection reward function optimization model, the problems of uneven distribution and high operating costs in bank branch site selection were solved, achieving more efficient and accurate branch site selection and flexible operations.

CN120672385APending Publication Date: 2025-09-19INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511108303.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies lack dynamic adaptability in bank branch site selection, resulting in uneven branch distribution, duplicate construction in some areas, and service gaps in remote areas. Furthermore, operating costs are high, and site selection decisions are not scientific and accurate.

Method used

A reinforcement deep learning-based method is adopted to obtain multi-dimensional dynamic joint description data, use the pre-trained reinforcement deep learning dynamic site selection model, combine the site selection reward function to optimize the model parameters, generate the target selection address, and provide feedback to the user to complete the dynamic site selection.

Benefits of technology

It improves the efficiency and accuracy of bank branch site selection, reduces operating costs, enhances the flexibility and dynamic adaptability of site selection, and optimizes branch distribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672385A_ABST
    Figure CN120672385A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic site selection method and device based on enhanced deep learning and a program product, and relates to the field of artificial intelligence. The method comprises the steps of obtaining current multi-dimensional dynamic joint description data corresponding to a target dynamic site selection area; inputting the current user distribution description data and the current website condition description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address; and the target selection address is fed back to the user, so that the user can complete dynamic site selection of the target dynamic site selection area according to the target selection address. The problems of low site selection efficiency, low coverage efficiency, high operation cost, poor dynamic adaptability and the like in the traditional bank outlet site selection process are solved, the bank outlet site selection efficiency and accuracy are improved, the address is selected according to the selected target, the operation cost is reduced, and the site selection flexibility is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a dynamic site selection method, device, and program product based on enhanced deep learning. Background Art

[0002] Based on people's demand for banks in daily life, it is important to reasonably increase the number of bank branches. However, for the selection of bank branch addresses, existing technologies generally use human experience judgment, statistical analysis or geographic information system analysis.

[0003] In the process of realizing the present invention, the inventors found that the existing technology has the following defects: At present, for judgment based on human experience, it lacks dynamic adaptability; for statistical analysis, it is impossible to accurately calculate the coverage effect of different site selection schemes on users, which easily leads to uneven distribution of network points, repeated construction in some areas, and lack of services in remote areas. The labor cost is high and the optimal allocation of resources cannot be achieved; for geographic information system analysis, it relies too much on human experience or simple data statistics, lacks systematic analysis of the interaction of complex factors, and is difficult to quantitatively evaluate the advantages and disadvantages of site selection schemes, resulting in poor scientificity and accuracy of site selection decisions. Summary of the Invention

[0004] The present invention provides a dynamic site selection method, device and program product based on enhanced deep learning to improve the efficiency and accuracy of bank branch site selection.

[0005] According to one aspect of the present invention, a dynamic site selection method based on reinforcement deep learning is provided, which includes:

[0006] Obtain the current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area;

[0007] The current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data;

[0008] Input the target dynamic site selection area, the current user distribution description data, and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address;

[0009] The enhanced deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data;

[0010] The target selection address is fed back to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

[0011] According to another aspect of the present invention, a dynamic site selection device based on enhanced deep learning is provided, comprising:

[0012] The current multi-dimensional dynamic joint description data acquisition module is used to obtain the current multi-dimensional dynamic joint description data corresponding to the target dynamic location area;

[0013] The current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data;

[0014] A target selection address generation module is used to input the target dynamic site selection area, the current user distribution description data and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address;

[0015] The enhanced deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data;

[0016] The target selection address feedback module is used to feed back the target selection address to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

[0017] According to another aspect of the present invention, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the dynamic site selection method based on reinforcement deep learning as described in any embodiment of the present invention is implemented.

[0018] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the dynamic site selection method based on reinforcement deep learning described in any embodiment of the present invention when executed.

[0019] According to another aspect of the present invention, a computer program product is provided, which includes a computer program, and when the computer program is executed by a processor, it implements the dynamic site selection method based on reinforcement deep learning described in any embodiment of the present invention.

[0020] The technical solution of the embodiment of the present invention obtains current multi-dimensional dynamic joint description data corresponding to a target dynamic location selection area; wherein the current multi-dimensional dynamic joint data includes current user distribution description data and current branch status description data; inputs the target dynamic location selection area, the current user distribution description data, and the current branch status description data into a pre-trained reinforcement deep learning dynamic location selection model to generate a target selection address; wherein the reinforcement deep learning dynamic location selection model optimizes model parameters based on a location selection reward function; the location selection reward function is designed based on the user distribution description data and the branch status description data; and the target selection address is fed back to the user to enable the user to complete the dynamic location selection of the target dynamic location selection area based on the target selection address. The technical solution solves the problems of low location selection efficiency, low coverage efficiency, high operating costs, and poor dynamic adaptability in the traditional bank branch location selection process, improves the efficiency and accuracy of bank branch location selection, reduces operating costs based on the selected target selection address, and improves location selection flexibility.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 is a flow chart of a dynamic site selection method based on enhanced deep learning provided according to embodiment 1 of the present invention;

[0024] Figure 2 is a detailed flow chart of a dynamic site selection method based on enhanced deep learning provided according to the second embodiment of the present invention;

[0025] Figure 3 2 is a schematic structural diagram of a dynamic site selection device based on enhanced deep learning according to the third embodiment of the present invention;

[0026] Figure 4 It is a structural diagram of an electronic device provided according to the fourth embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "target", "current", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0029] It is worth noting that in the technical solution of this application, the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse; if the user chooses to refuse, the expert decision-making process will be entered.

[0030] Example 1

[0031] Figure 1 A flowchart of a dynamic site selection method based on reinforcement deep learning is provided for the first embodiment of the present invention. This embodiment is applicable to the situation where a target website is selected for a bank. The method can be executed by a dynamic site selection device based on reinforcement deep learning, and the dynamic site selection device based on reinforcement deep learning can be implemented in the form of hardware and / or software.

[0032] Correspondingly, such as Figure 1 As shown, the method includes:

[0033] S110: Obtain current multi-dimensional dynamic joint description data corresponding to the target dynamic location area.

[0034] The current multi-dimensional dynamic combined data includes current user distribution description data and current network status description data.

[0035] The target dynamic location selection area may be an area corresponding to a specific location range determined based on user needs. The current multi-dimensional dynamic joint description data may be data describing user distribution and network conditions.

[0036] For example, suppose a user wants to determine the location of a bank branch in a newly constructed town area A. Town area A can be abstracted as a two-dimensional space with length L and width W. The local branch plans to build m new branches to serve the area's n permanent residents. The goal is to maximize the service revenue provided by these m branches, that is, to maximize the location reward function.

[0037] S120: Input the target dynamic site selection area, the current user distribution description data, and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address.

[0038] Among them, the reinforced deep learning dynamic site selection model optimizes the model parameters based on the site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data.

[0039] Among them, the enhanced deep learning dynamic site selection model can be a model that can target address selection in a certain area by combining user and network parameters.

[0040] Continuing with the previous example, we can factor in the service capacity of outlets to reduce the location reward function (simulating expenditure) and increase it as more users are served (simulating revenue generated by user behavior). Ultimately, through model training, the location reward function converges to a theoretically optimal value. The selected outlet locations and the expected service capacity at this point will serve as a reference for site selection in Town Area A.

[0041] Optionally, the target dynamic site selection area, the current user distribution description data and the current network status description data are input into a pre-trained reinforcement deep learning dynamic site selection model to generate a target selection address, including: obtaining the geographical locations of each current network point corresponding to the target dynamic site selection area, and generating current observation information; inputting the current user distribution description data and the current network status description data into a pre-trained reinforcement deep learning dynamic site selection model to generate at least one current selection address and a site selection service capability corresponding to each current selection address; and selecting the target selection address based on the current observation information and the site selection service capability, combined with the site selection reward function in the reinforcement deep learning dynamic site selection model.

[0042] Among them, the current observation information can be determined and processed through the geographical location of each current branch corresponding to the target dynamic site selection area. Specifically, the geographical information of multiple bank branches is determined and used as observation information to realize the selection of new bank branch addresses.

[0043] For example, assuming there are m bank branches, the current observation information is determined based on the set of geographic locations. At the beginning, the model will randomly generate m geographic locations and initial service capabilities, calculate the reward function based on the current observation information, and continuously optimize the location information and service capacity increase and decrease through the reinforcement deep learning dynamic site selection model, ultimately making the branch location converge to an optimal solution set, that is, determining the target selection address.

[0044] The advantage of this setting is that the current observation information is determined in combination with the current network point geographical location, and then the target selection address is determined. In this way, the target selection address can be determined more accurately through the determination of the observation information.

[0045] S130: Feedback the target selected address to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selected address.

[0046] In this embodiment, after the target selection address is determined, the target selection address is fed back to the user, and the user can use the target selection address to determine the bank website in the target dynamic address selection area. This can save labor costs and the selected target selection address is more in line with the requirements and more accurate.

[0047] The technical solution of the embodiment of the present invention obtains current multi-dimensional dynamic joint description data corresponding to a target dynamic location selection area; wherein the current multi-dimensional dynamic joint data includes current user distribution description data and current branch status description data; inputs the target dynamic location selection area, the current user distribution description data, and the current branch status description data into a pre-trained reinforcement deep learning dynamic location selection model to generate a target selection address; wherein the reinforcement deep learning dynamic location selection model optimizes model parameters based on a location selection reward function; the location selection reward function is designed based on the user distribution description data and the branch status description data; and the target selection address is fed back to the user to enable the user to complete the dynamic location selection of the target dynamic location selection area based on the target selection address. The technical solution solves the problems of low location selection efficiency, low coverage efficiency, high operating costs, and poor dynamic adaptability in the traditional bank branch location selection process, improves the efficiency and accuracy of bank branch location selection, reduces operating costs based on the selected target selection address, and improves location selection flexibility.

[0048] Example 2

[0049] Figure 2 This is a detailed flowchart of a dynamic site selection method based on enhanced deep learning, provided according to Example 2 of the present invention. This embodiment is optimized based on the above embodiments. In this embodiment, before obtaining the current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area, it also includes training the enhanced deep learning dynamic site selection model.

[0050] S210: Obtain each historical dynamic site selection area.

[0051] S220 , sequentially obtaining a target historical dynamic location selection area and historical multi-dimensional dynamic joint data corresponding to each of the target historical dynamic location selection areas.

[0052] The historical multi-dimensional dynamic joint data includes historical user distribution description data and historical network status description data.

[0053] S230 . Divide the target historical dynamic location selection area by a preset regional grid division method to obtain the target historical dynamic location selection grid division area.

[0054] The regional grid division method may be a method of dividing the historical dynamic location selection area, which can better determine the bank's geographical location. The target historical dynamic location selection grid division area may be a division area result obtained after dividing the historical dynamic location selection area.

[0055] For example, each target historical dynamic site selection area can be divided according to preset division parameters. Specifically, it can be divided into 200×200 grid units, and each grid unit is 250m×250m in size, so as to analyze various indicators in the area in detail.

[0056] S240: Obtain and form a historical network point geographical location set based on the geographical location of each historical network point, and generate current historical observation information.

[0057] S250. Divide the area according to the current historical observation information and the target historical dynamic site selection grid, train the initial reinforcement deep learning dynamic site selection model through the historical user distribution description data and the historical network status description data, and optimize the parameters of the trained initial reinforcement deep learning dynamic site selection model in combination with the pre-designed site selection reward function to determine whether the preset training completion parameter threshold requirements are met. If not, return to execute the operation of obtaining a target historical dynamic site selection area in turn until the training completion parameter threshold requirements are met, and obtain the trained reinforcement deep learning dynamic site selection model.

[0058] Optionally, the site selection reward function includes a positive benefit reward part and a negative benefit reward part; wherein, the positive benefit reward part is used to represent the size of the site selection service capability; the negative benefit reward part is used to represent the size of the cost required for the outlet to provide the site selection service capability; according to the historical user distribution description data and the historical outlet status description data, the positive benefit reward part and the negative benefit reward part are calculated respectively; the negative benefit reward part is subtracted from the positive benefit reward part to obtain the site selection reward function.

[0059] For example, assume that the site selection reward function is R and the positive income reward part is R U ; The negative income reward is R B ; For R U Calculated from historical user distribution description data, R B Calculated from historical site status description data. Among them, the site selection reward function can be calculated as R = R U -R B .

[0060] The advantage of this setting is that: through the location reward function consisting of a positive benefit reward part and a negative benefit reward part, as well as the calculation process of the location reward function, a better location reward function can be obtained, so that the trained reinforced deep learning dynamic location selection model is more accurate, and further can assist in more accurately determining the target selection address.

[0061] Optionally, the positive profit reward part and the negative profit reward part are calculated separately based on the historical user distribution description data and the historical network status description data, including: wherein the historical user distribution description data includes multiple users, and a historical user set is determined based on each of the users; the historical network status description data includes multiple networks, and a historical network set is determined based on each of the networks; in the historical user set, each user is traversed in turn, and the positive profit reward part is calculated based on the exponential function of whether each user is served by the network; in the historical network set, each network is traversed in turn, and the negative profit reward part is obtained according to the preset negative profit reward calculation method.

[0062] For example, it is assumed that the historical user set is U, which includes n users u; the historical network point set is B, which includes m networks b.

[0063] Since in the historical user set, each user is traversed in turn, and the positive income reward part is calculated according to the exponential function of whether each user is served by the outlet, it can be known that the positive income reward part When the user is served, the value of the parameter can be set to 1; when the user is not served, the value of the parameter can be set to 0.

[0064] Specifically, the condition for user u to be served is that: it is located within the coverage area of ​​at least one network point b (the coverage area of ​​the circular area with network point b as the center and a preset length r as the radius) and the service capability c of network point b is b Not exhausted; when calculating the positive income reward part, for each user u, check its distance to all outlets. If there is an outlet b that satisfies distance(u,b)≤r and the outlet still has remaining service capacity, the user is served, that is, Otherwise, the conditions are not met.

[0065] The advantage of this setting is that by combining multiple users and multiple outlets in the historical user distribution description data and the historical outlet status description data to calculate the positive benefit reward part and the negative benefit reward part respectively, the positive and negative benefit reward parts can be considered and calculated from the user and outlet conditions in multiple dimensions to construct a more accurate site selection reward function.

[0066] Optionally, each outlet corresponds to the historical outlet service capability, service capability cost coefficient, historical outlet location user density and outlet location rental cost coefficient; the negative benefit reward part is obtained according to the preset negative benefit reward calculation method, including: according to the preset negative benefit reward calculation method, multiplying the outlet service capability and the service capability cost coefficient to obtain the weighted outlet service capability; multiplying the historical outlet location user density and the outlet location rental cost coefficient to obtain the weighted outlet rental cost; adding the weighted outlet service capability and the weighted outlet rental cost to obtain the negative benefit reward part.

[0067] For example, assuming that the historical service capacity of the network is c b , the service capacity cost coefficient is α, and the user density of the historical network location is ρ b , the rental cost coefficient of the location of the outlet is β. Then we can calculate the negative income reward part as R B =∑ b∈B α*c b +β*ρ b .

[0068] Among them, the weighted network service capacity is composed of α*c b Calculated, α reflects variable costs such as employee salaries, which require calibration based on actual costs. The location cost coefficient for a branch's location is positively correlated with the user density of the historical branch location, meaning that rents are higher in densely populated areas. This requires calibration based on actual market rents.

[0069] Furthermore, we can calculate By setting it up in this way, we can determine the target selection address that covers high-density user areas, avoids excessive rent, meets demand but is not redundant.

[0070] The advantage of this setting is that the negative profit reward part is calculated by using the historical outlet service capacity, service capacity cost coefficient, historical outlet location user density and outlet location rental cost coefficient corresponding to each outlet. It can consider various situations of the outlets from multiple dimensions and calculate a more accurate negative profit reward part, which is conducive to the calculation of the site selection reward function.

[0071] S260: Obtain current multi-dimensional dynamic joint description data corresponding to the target dynamic location area.

[0072] The current multi-dimensional dynamic combined data includes current user distribution description data and current network status description data.

[0073] S270: Input the target dynamic site selection area, the current user distribution description data, and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address.

[0074] S280: Feedback the target selection address to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

[0075] The technical solution of the embodiment of the present invention is that the location reward function is composed of a positive benefit reward part and a negative benefit reward part, and the location reward function is calculated in a process, so that a better location reward function can be obtained, so that the trained reinforcement deep learning dynamic location selection model is more accurate, and further can assist in more accurately determining the target selection address; by combining multiple users and multiple outlets in the historical user distribution description data and the historical outlet status description data to respectively calculate the positive benefit reward part and the negative benefit reward part, the positive and negative benefit reward parts can be considered and calculated from the user and outlet conditions in multiple dimensions, so as to construct a more accurate location reward function. function; the negative benefit reward part is calculated by the historical outlet service capacity, service capacity cost coefficient, historical outlet location user density and outlet location rental cost coefficient corresponding to each outlet. It can consider various situations of the outlets in multiple dimensions and calculate a more accurate negative benefit reward part, which is conducive to the calculation of the site selection reward function; through the training determination process of the enhanced deep learning dynamic site selection model, a more accurate enhanced deep learning dynamic site selection model is trained, so that the target selection address can be determined more accurately. The site is selected by the enhanced deep learning dynamic site selection model, which reduces labor costs and improves the efficiency and flexibility of site selection.

[0076] Example 3

[0077] Figure 3 This is a structural diagram of a dynamic site selection device based on enhanced deep learning provided in the third embodiment of the present invention. The dynamic site selection device based on enhanced deep learning provided in this embodiment can be implemented through software and / or hardware, and can be configured in a terminal device or server to implement a dynamic site selection method based on enhanced deep learning in the embodiment of the present invention. Figure 3 As shown, the device includes: a current multi-dimensional dynamic joint description data acquisition module 310, a target selection address generation module 320 and a target selection address feedback module 330.

[0078] The current multi-dimensional dynamic joint description data acquisition module 310 is used to acquire the current multi-dimensional dynamic joint description data corresponding to the target dynamic location area;

[0079] The current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data;

[0080] The target selection address generation module 320 is used to input the target dynamic site selection area, the current user distribution description data and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address;

[0081] The enhanced deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data;

[0082] The target selection address feedback module 330 is used to feed back the target selection address to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

[0083] The technical solution of the embodiment of the present invention obtains current multi-dimensional dynamic joint description data corresponding to a target dynamic location selection area; wherein the current multi-dimensional dynamic joint data includes current user distribution description data and current branch status description data; inputs the target dynamic location selection area, the current user distribution description data, and the current branch status description data into a pre-trained reinforcement deep learning dynamic location selection model to generate a target selection address; wherein the reinforcement deep learning dynamic location selection model optimizes model parameters based on a location selection reward function; the location selection reward function is designed based on the user distribution description data and the branch status description data; and the target selection address is fed back to the user to enable the user to complete the dynamic location selection of the target dynamic location selection area based on the target selection address. The technical solution solves the problems of low location selection efficiency, low coverage efficiency, high operating costs, and poor dynamic adaptability in the traditional bank branch location selection process, improves the efficiency and accuracy of bank branch location selection, reduces operating costs based on the selected target selection address, and improves location selection flexibility.

[0084] Based on the above embodiments, the target selection address generation module can be specifically used to: obtain the geographical location of each current outlet corresponding to the target dynamic site selection area and generate current observation information; input the current user distribution description data and the current outlet status description data into a pre-trained reinforcement deep learning dynamic site selection model to generate at least one current selection address and the site selection service capability corresponding to each current selection address; based on the current observation information and site selection service capability, combined with the site selection reward function in the reinforcement deep learning dynamic site selection model, select the target selection address.

[0085] On the basis of the above embodiments, it also includes a strengthened deep learning dynamic site selection model training module, which is used to obtain each historical dynamic site selection area before obtaining the current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area; sequentially obtain a target historical dynamic site selection area, and historical multi-dimensional dynamic joint data corresponding to the target historical dynamic site selection area; wherein the historical multi-dimensional dynamic joint data includes historical user distribution description data and historical network status description data; divide the target historical dynamic site selection area by a preset regional grid division method to obtain the target historical dynamic site selection grid division area; obtain and calculate the target historical dynamic site selection grid division area according to the geographical location of each historical network point; , forming a set of historical network location locations, and generating current historical observation information; dividing the area according to the current historical observation information and the target historical dynamic site selection grid, training the initial reinforcement deep learning dynamic site selection model through the historical user distribution description data and the historical network status description data, and combining the pre-designed site selection reward function to optimize the parameters of the trained initial reinforcement deep learning dynamic site selection model, and judge whether the preset training completion parameter threshold requirement is met. If not, return to execute the operation of obtaining a target historical dynamic site selection area in turn until the training completion parameter threshold requirement is met, and obtain the trained reinforcement deep learning dynamic site selection model.

[0086] Based on the above embodiments, the site selection reward function includes a positive profit reward part and a negative profit reward part; wherein, the positive profit reward part is used to represent the size of the site selection service capability; the negative profit reward part is used to represent the size of the cost required for the outlet to provide the site selection service capability.

[0087] Based on the above embodiments, the enhanced deep learning dynamic site selection model training module may further specifically include: a positive benefit reward part and a negative benefit reward part calculation unit, which is used to calculate the positive benefit reward part and the negative benefit reward part respectively according to the historical user distribution description data and the historical network status description data; a site selection reward function determination unit, which is used to subtract the negative benefit reward part from the positive benefit reward part to obtain the site selection reward function.

[0088] On the basis of the above embodiments, the positive profit reward part and the negative profit reward part calculation unit can be specifically used for: wherein, the historical user distribution description data includes multiple users, and the historical user set is determined based on each of the users; the historical network status description data includes multiple networks, and the historical network set is determined based on each of the networks; the positive profit reward part calculation subunit is used to traverse each user in the historical user set in turn, and calculate the positive profit reward part according to the exponential function of whether each user is served by the network; the negative profit reward part calculation subunit is used to traverse each network in the historical network set in turn, and obtain the negative profit reward part according to the preset negative profit reward calculation method.

[0089] On the basis of the above embodiments, each network point corresponds to the historical network point service capacity, service capacity cost coefficient, historical network point location user density and network point location rental cost coefficient.

[0090] On the basis of the above embodiments, the negative profit reward part calculation sub-unit can be specifically used to: multiply the outlet service capacity and the service capacity cost coefficient according to the preset negative profit reward calculation method to obtain the weighted outlet service capacity; multiply the user density of the historical outlet location and the rental cost coefficient of the outlet location to obtain the weighted outlet rental cost; add the weighted outlet service capacity and the weighted outlet rental cost to obtain the negative profit reward part.

[0091] The dynamic site selection device based on reinforcement deep learning provided in an embodiment of the present invention can execute the dynamic site selection method based on reinforcement deep learning provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0092] Example 4

[0093] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement the fourth embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0094] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0095] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0096] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the dynamic site selection method based on reinforcement deep learning.

[0097] In some embodiments, the dynamic site selection method based on reinforcement deep learning can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the dynamic site selection method based on reinforcement deep learning described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to execute the dynamic site selection method based on reinforcement deep learning by any other appropriate means (for example, by means of firmware).

[0098] The method includes: obtaining current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area; wherein the current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data; inputting the target dynamic site selection area, the current user distribution description data and the current network status description data into a pre-trained reinforcement deep learning dynamic site selection model to generate a target selection address; wherein the reinforcement deep learning dynamic site selection model is based on a site selection reward function to optimize the model parameters; the site selection reward function is designed based on the user distribution description data and the network status description data; the target selection address is fed back to the user to enable the user to complete the dynamic site selection of the target dynamic site selection area based on the target selection address.

[0099] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0100] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0101] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0102] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0103] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0104] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0105] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0106] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

[0107] Example 5

[0108] Embodiment 5 of the present invention also provides a computer-readable storage medium, wherein the computer-readable instructions, when executed by a computer processor, are used to execute a dynamic site selection method based on reinforcement deep learning, the method comprising: obtaining current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area; wherein the current multi-dimensional dynamic joint data comprises current user distribution description data and current network status description data; inputting the target dynamic site selection area, the current user distribution description data and the current network status description data into a pre-trained reinforcement deep learning dynamic site selection model to generate a target selection address; wherein the reinforcement deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data; and feeding back the target selection address to the user to enable the user to complete the dynamic site selection of the target dynamic site selection area based on the target selection address.

[0109] Of course, an embodiment of the present invention provides a computer-readable storage medium, and its computer-executable instructions are not limited to the method operations described above, but can also execute related operations in dynamic site selection based on reinforcement deep learning provided by any embodiment of the present invention.

[0110] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general-purpose hardware, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.

[0111] It is worth noting that in the above-mentioned embodiment of dynamic site selection based on enhanced deep learning, the various units and modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0112] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A dynamic site selection method based on reinforcement deep learning, characterized in that: include: Obtain the current multi-dimensional dynamic joint description data corresponding to the target dynamic site selection area; The current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data; Input the target dynamic site selection area, the current user distribution description data, and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address; The enhanced deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data; The target selection address is fed back to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

2. The method according to claim 1, characterized in that The step of inputting the target dynamic site selection area, the current user distribution description data, and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address includes: Obtain the geographical location of each current network point corresponding to the target dynamic site selection area and generate current observation information; Inputting the current user distribution description data and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate at least one currently selected address and the site selection service capability corresponding to each currently selected address; According to the current observation information and site selection service capabilities, combined with the site selection reward function in the enhanced deep learning dynamic site selection model, the target selection address is selected.

3. The method according to claim 1, characterized in that Before obtaining the current multi-dimensional dynamic joint description data corresponding to the target dynamic location area, the method further includes: Obtain each historical dynamic site selection area; Sequentially acquiring a target historical dynamic location area and historical multi-dimensional dynamic joint data corresponding to each of the target historical dynamic location areas; The historical multi-dimensional dynamic joint data includes historical user distribution description data and historical network status description data; Dividing the target historical dynamic location area by a preset regional grid division method to obtain the target historical dynamic location grid division area; Obtain and form a historical network point geographical location set based on the geographical location of each historical network point, and generate current historical observation information; According to the current historical observation information and the target historical dynamic site selection grid division area, the initial reinforcement deep learning dynamic site selection model is trained through the historical user distribution description data and the historical network status description data, and combined with the pre-designed site selection reward function, the parameters of the trained initial reinforcement deep learning dynamic site selection model are optimized to determine whether the preset training completion parameter threshold requirements are met. If not, the operation of obtaining a target historical dynamic site selection area in turn is returned to execute until the training completion parameter threshold requirements are met, thereby obtaining a trained reinforcement deep learning dynamic site selection model.

4. The method according to claim 3, characterized in that The site selection reward function includes a positive benefit reward part and a negative benefit reward part; Among them, the positive benefit reward part is used to represent the size of the site selection service capability; the negative benefit reward part is used to represent the size of the cost required for the outlet to provide the site selection service capability; Based on historical user distribution description data and historical network status description data, the positive income reward part and the negative income reward part are calculated respectively; The location reward function is obtained by subtracting the negative benefit reward part from the positive benefit reward part.

5. The method according to claim 4, characterized in that Based on the historical user distribution description data and the historical network status description data, the positive income reward part and the negative income reward part are calculated respectively. include: The historical user distribution description data includes multiple users, and the historical user set is determined based on each of the users; the historical network status description data includes multiple networks, and the historical network set is determined based on each of the networks; In the historical user set, each user is traversed in turn, and the positive income reward portion is calculated based on the exponential function of whether each user is served by the outlet; In the historical network set, each network is traversed in turn, and the negative profit reward part is obtained according to the preset negative profit reward calculation method.

6. The method according to claim 5, characterized in that The historical service capacity of each outlet, the service capacity cost coefficient, the user density of the historical outlet location, and the rental cost coefficient of the outlet location; The negative income reward part is obtained according to the preset negative income reward calculation method, including: According to a preset negative income reward calculation method, the outlet service capacity and the service capacity cost coefficient are multiplied to obtain a weighted outlet service capacity; Multiplying the user density at the location of the historical network point and the rental cost coefficient at the location of the network point to obtain a weighted network point rental cost; The weighted outlet service capacity and the weighted outlet rental cost are added together to obtain the negative income reward part.

7. A dynamic site selection device based on enhanced deep learning, characterized in that: include: The current multi-dimensional dynamic joint description data acquisition module is used to obtain the current multi-dimensional dynamic joint description data corresponding to the target dynamic location area; The current multi-dimensional dynamic joint data includes current user distribution description data and current network status description data; A target selection address generation module is used to input the target dynamic site selection area, the current user distribution description data and the current network status description data into a pre-trained enhanced deep learning dynamic site selection model to generate a target selection address; The enhanced deep learning dynamic site selection model optimizes model parameters based on a site selection reward function; the site selection reward function is designed based on user distribution description data and network status description data; The target selection address feedback module is used to feed back the target selection address to the user, so that the user can complete the dynamic location selection of the target dynamic location selection area according to the target selection address.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, it implements a dynamic site selection method based on reinforcement deep learning as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a dynamic site selection method based on reinforcement deep learning as described in any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the dynamic site selection method based on reinforcement deep learning according to any one of claims 1 to 6.

Citation Information

Cited By

  • Chain store site selection optimization method and system based on reinforcement learning

    CN121860694A