Store data processing method and device

By using semantic similarity algorithms and multiple artificial intelligence models in a collaborative manner, the problem of low efficiency in processing store data under the manual annotation mode has been solved, realizing automated and accurate updates of store data and improving the timeliness and accuracy of the data.

CN121807834APending Publication Date: 2026-04-07ALIPAY COM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing technologies, manual annotation is inefficient and cannot meet the batch processing needs of large-scale store data, resulting in a long data preservation period and affecting the timeliness and accuracy of data at the decision-making level.

Method used

By using semantic similarity algorithms to identify similar stores of the target store from online service platforms or offline databases, and by using multiple artificial intelligence models to make collaborative judgments, the data of the same store can be accurately identified and processed, reducing the reliance on manual intervention.

Benefits of technology

It enables automated and precise updates of store data, improves the timeliness and accuracy of data, reduces reliance on manual labor, and adapts to the processing needs of large-scale store data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121807834A_ABST
    Figure CN121807834A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a store data processing method and device. First store data of a target store is obtained from an offline database; then, acquiring a similar store set corresponding to the target store from an online service platform or an offline database; under the condition that the similar store set comprises at least one similar store, the semantic similarity between the target store and the target similar store is determined, and under the condition that the semantic similarity is smaller than a first similarity threshold value, whether the first store data and the second store data of the target similar store belong to the same store or not is recognized through multiple artificial intelligence models; and when the plurality of artificial intelligence models are utilized to identify that the first store data and the second store data belong to the same store, or the semantic similarity is greater than or equal to a first similarity threshold, performing data processing on the first store data in the offline database based on the second store data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of computer network, and particularly relate to a store data processing method and device. BACKGROUND

[0002] In the field of local life services, retail chain operations and the like that rely on store data to make operational decisions, the timeliness and accuracy of store data affect the execution effect of the downstream decision-making end. Therefore, a data preservation processing mechanism can be used to periodically maintain the store data stored in the offline database to ensure that the offline data is synchronized with the actual operating status of the store.

[0003] In related technologies, data preservation processing is mainly in the mode of manual annotation, that is, manually verifying, entering and updating store data in the offline database. However, as the scale of the online service platform expands, the number of stores entering the platform increases explosively, and the update frequency of store information also continues to increase. The manual annotation mode has low processing efficiency and is difficult to match the batch processing needs of large-scale store data, resulting in a long data preservation period.

[0004] Therefore, there is an urgent need for a more efficient data preservation processing scheme to meet the processing needs of large-scale store data, reduce the degree of dependence on manual work, and provide reliable data support for the decision-making end. SUMMARY

[0005] To meet the processing needs of large-scale store data and reduce the degree of dependence on manual work, one or more embodiments of the present specification provide a store data processing method and device.

[0006] In a first aspect, one or more embodiments of the present specification provide a store data processing method, which comprises: obtaining first store data of a target store from an offline database; obtaining a set of similar stores corresponding to the target store; in a case where the set of similar stores includes at least one similar store, determining a semantic similarity between the target store and the target similar store based on the first store data and second store data, wherein the second store data is store data corresponding to a target similar store in the set of similar stores; in a case where the semantic similarity is less than a first similarity threshold, using multiple artificial intelligence models to identify whether the first store data and the second store data belong to the same store; in a case where the first store data and the second store data are identified as belonging to the same store by using the multiple artificial intelligence models, or the semantic similarity is greater than or equal to the first similarity threshold, performing data processing on the first store data in the offline database based on the second store data.

[0007] In a possible implementation, the acquiring the set of similar stores corresponding to the target store comprises: acquiring longitude and latitude information corresponding to each transaction data of the target store in a preset time period; performing clustering processing on the longitude and latitude information; in a case where a target longitude and latitude that can represent a location of the target store exists in a clustering result, acquiring, according to the target longitude and latitude and a store name of the target store, the set of similar stores corresponding to the target store from an online service platform; in a case where the target longitude and latitude that can represent the location of the target store does not exist in the clustering result, acquiring, according to the store name and a store address of the target store, the set of similar stores corresponding to the target store from the online service platform; and the second store data is store data of a target similar store in the set of similar stores on the online service platform.

[0008] In a possible implementation, in a case where a target longitude and latitude that can represent a location of the target store exists in a clustering result, the acquiring the set of similar stores corresponding to the target store according to the target longitude and latitude and a store name of the target store from an online service platform comprises: in the case where the target longitude and latitude that can represent the location of the target store exists in the clustering result, calling an interface of the online service platform, and inputting a retrieval parameter to the interface, the retrieval parameter comprising a preset geographical range centered on the target longitude and latitude and a fuzzy matching rule of the store name; and receiving a set of similar stores that satisfy the retrieval parameter and that are fed back by the interface.

[0009] In a possible implementation, in a case where the first store data and the second store data belong to the same store, or the semantic similarity is greater than or equal to the first similarity threshold, by using the plurality of artificial intelligence models, the data processing on the first store data in the offline database based on the second store data comprises: in the case where the first store data and the second store data belong to the same store, or the semantic similarity is greater than or equal to the first similarity threshold, by using the plurality of artificial intelligence models, updating the first store data in the offline database based on the second store data.

[0010] In a possible implementation, the acquiring the set of similar stores corresponding to the target store comprises: determining store similarities respectively corresponding to the target store and other stores in the offline database; and acquiring, from the offline database, stores that belong to a same service scenario as the target store and have a store similarity greater than a second similarity threshold, to obtain the set of similar stores corresponding to the target store; and the second store data is store data of a target similar store in the set of similar stores in the offline database.

[0011] In one possible implementation, if multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then data processing is performed on the first store data in the offline database based on the second store data. This includes: deduplicating the first store data and the second store data in the offline database if multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold.

[0012] In one possible implementation, obtaining the first store data of the target store from an offline database includes: receiving a selection operation on data tags, wherein different data tags correspond to different store filtering conditions; determining a set of stores to be processed based on the store filtering conditions corresponding to the selected data tags, wherein the set of stores to be processed includes the target store; and obtaining the first store data of the target store from the offline database.

[0013] In one possible implementation, the first store data and the second store data include store data across multiple dimensions. Where the set of similar stores includes at least one similar store, determining the semantic similarity between the target store and the target similar store based on the first store data and the second store data includes: when the set of similar stores includes at least one similar store, calculating the single-dimensional similarity between the first store data and the second store data for each dimension of the store data; and weighting the single-dimensional similarities for each dimension based on their respective weight values ​​to obtain the semantic similarity between the target store and the target similar store.

[0014] In one possible implementation, multiple artificial intelligence models are used to identify whether the first store data and the second store data belong to the same store. This includes: inputting the first store data, the second store data, and prompt words into multiple large language models respectively, wherein the prompt words are used to instruct the large language models to identify whether the two stores are the same store. Based on the recognition results of various language models and preset voting rules, it is determined whether the target store and the target similar store are the same store.

[0015] In one possible implementation, determining whether the target store and the target similar store are the same store based on the recognition results of various language models and preset voting rules includes: counting the number of target recognition results in the recognition results of various language models, wherein the target recognition result is the recognition result that indicates that the target store and the target similar store are the same store; calculating the proportion of the number of target recognition results to the total number of recognition results; and determining that the target store and the target similar store are the same store if the proportion exceeds a preset proportion threshold.

[0016] Secondly, one or more embodiments of this specification also provide a store data processing apparatus, the apparatus comprising: a first acquisition module, configured to acquire first store data of a target store from an offline database; a second acquisition module, configured to acquire a set of similar stores corresponding to the target store; a determination module, configured to determine the semantic similarity between the target store and the target similar store based on the first store data and the second store data, wherein the second store data is the store data corresponding to the target similar store in the set of similar stores; an identification module, configured to identify whether the first store data and the second store data belong to the same store using multiple artificial intelligence models when the semantic similarity is less than a first similarity threshold; and a data processing module, configured to perform data processing on the first store data in the offline database based on the second store data when the first store data and the second store data are identified as belonging to the same store using multiple artificial intelligence models, or when the semantic similarity is greater than or equal to the first similarity threshold.

[0017] In one possible implementation, the second acquisition module is configured to: acquire the latitude and longitude information corresponding to each transaction data of the target store within a preset time period; perform clustering processing on the latitude and longitude information; if there is a target latitude and longitude that can characterize the location of the target store in the clustering results, acquire a set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store; if there is no target latitude and longitude that can characterize the location of the target store in the clustering results, acquire a set of similar stores corresponding to the target store from the online service platform according to the store name and store address of the target store.

[0018] In one possible implementation, the second acquisition module is used to, when a target latitude and longitude that can characterize the location of the target store exists in the clustering results, obtain a set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store. Specifically, when a target latitude and longitude that can characterize the location of the target store exists in the clustering results, the module calls the interface of the online service platform and passes search parameters to the interface. The search parameters include a preset geographical range centered on the target latitude and longitude and a fuzzy matching rule for the store name. The module receives the set of similar stores that satisfy the search parameters from the interface. The second store data is the store data of the target similar stores in the set of similar stores on the online service platform.

[0019] In one possible implementation, the data processing module is configured to: update the first store data in the offline database based on the second store data, provided that the first store data and the second store data belong to the same store, or the semantic similarity is greater than or equal to a first similarity threshold, using multiple artificial intelligence models.

[0020] In one possible implementation, the second acquisition module is configured to: determine the store similarity between the target store and other stores in the offline database; acquire stores from the offline database that belong to the same service scenario as the target store and have a store similarity greater than a second similarity threshold, thereby obtaining a set of similar stores corresponding to the target store; wherein, the second store data is the store data of the target similar store in the offline database.

[0021] In one possible implementation, the data processing module is used to deduplicate the first store data and the second store data in the offline database when multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or when the semantic similarity is greater than or equal to a first similarity threshold.

[0022] In one possible implementation, the first acquisition module is configured to: receive a selection operation for data tags, wherein different data tags correspond to different store filtering conditions; determine a set of stores to be processed based on the store filtering conditions corresponding to the selected data tags, wherein the set of stores to be processed includes the target store; and acquire first store data of the target store from an offline database.

[0023] In one possible implementation, the first store data and the second store data include store data with multiple dimensions; wherein, the determining module is configured to: when the set of similar stores includes at least one similar store, calculate the single-dimensional similarity between the first store data and the second store data in the corresponding dimension for each dimension of the store data; and perform weighted processing on the single-dimensional similarity corresponding to each dimension based on the weight values ​​corresponding to each dimension to obtain the semantic similarity between the target store and the target similar store.

[0024] In one possible implementation, the recognition module is configured to: input the first store data, the second store data, and prompt words into multiple large language models respectively, wherein the prompt words are used to instruct the large language models to identify whether two stores are the same store; and determine whether the target store and the target similar store are the same store based on the recognition results of the large language models and preset voting rules.

[0025] In one possible implementation, the identification module is used to: count the number of target identification results among the identification results of various language models, wherein the target identification result is the identification result that indicates that the target store and the target similar store are the same store; calculate the proportion of the number of target identification results to the total number of identification results; and determine that the target store and the target similar store are the same store if the proportion exceeds a preset proportion threshold.

[0026] Thirdly, one or more embodiments of this specification also provide an electronic device, which includes a memory and a processor; the memory is used to store a computer program product; the processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, it implements the store data processing method of the first aspect described above.

[0027] Fourthly, one or more embodiments of this specification also provide a computer-readable storage medium storing computer program instructions that, when executed, implement the store data processing method of the first aspect described above.

[0028] In summary, one or more embodiments of this specification provide a method and apparatus for processing store data. The method first uses a semantic similarity algorithm to identify store data from an online service platform or offline database that belongs to the same entity as the target store in the offline database. For fuzzy matching scenarios where semantic similarity fails to identify the store, multiple artificial intelligence models can be introduced for collaborative judgment. Leveraging the deep analysis capabilities of AI models on multi-dimensional store data, the risk of misjudgment by a single identification method is avoided, accurately solving the problem of store identity identification in complex scenarios. Finally, after confirming that the stores belong to the same entity through semantic similarity verification or collaborative judgment by AI models, the corresponding store data in the offline database can be processed based on the identified store data, reducing reliance on manual intervention and achieving automated and accurate updates of store data, improving data timeliness and accuracy. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of one or more embodiments of this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of one or more embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This specification provides a schematic diagram of an application scenario for one or more embodiments. Figure 2 A flowchart illustrating a store data processing method provided in one or more embodiments of this specification; Figure 3 This specification provides a schematic diagram of a process for obtaining a set of similar stores from an online service platform, which is provided for one or more embodiments of this specification. Figure 4 This is a schematic diagram illustrating yet another process for obtaining a set of similar stores from an online service platform, provided for one or more embodiments of this specification. Figure 5 This specification provides a flowchart illustrating a process for retrieving a set of similar stores from an offline database, as illustrated in one or more embodiments of this specification. Figure 6 A structural block diagram of a store data processing device provided for one or more embodiments of this specification; Figure 7 This is a structural block diagram of an electronic device provided for one or more embodiments of this specification. Detailed Implementation

[0031] The following further details one or more embodiments of this specification through the accompanying drawings and examples. Through these descriptions, the features and advantages of one or more embodiments of this specification will become clearer and more definite.

[0032] The special term "exemplary" here means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" here does not necessarily have to be construed as superior or better than other embodiments. Although various aspects of the embodiments are shown in the accompanying drawings, unless specifically stated, the drawings do not have to be drawn to scale.

[0033] In addition, the technical features involved in different embodiments of one or more embodiments of this specification described below can be combined with each other as long as they do not conflict with each other.

[0034] For the convenience of understanding, the application scenarios of the technical solutions provided by one or more embodiments of this specification are described below first.

[0035] As Figure 1 shown, in fields such as local life services and retail chain operations that rely on store data support, there is a three-tier architecture relationship including an online service platform 01, an operation platform 02, and a decision-making end 03. Among them, the online service platform 01 refers to an open service platform facing users and merchants.线下实体门店可以通过在线服务平台01完成门店信息注册、展示及更新,例如,在线服务平台01可以是地图类应用 / 平台、外卖类应用 / 平台等。运营平台02是对在线服务平台01进行管理的后端运营支撑平台,其核心职责是整合在线服务平台上的门店数据,为决策端03提供数据支撑,三者形成门店信息上传至在线服务平台、运营平台同步整合数据、决策端基于数据开展运营决策的完整链路。

[0036] The offline database supporting the operation platform 02 is used to store store data. Store data usually covers multi-dimensional core information such as store name, address, longitude, and latitude. The offline database is the core data source for the decision-making end to carry out operation actions such as merchant management, precise recommendation, regional marketing, and service fulfillment. Its accuracy and timeliness affect the scientific nature of the decision-making end's analysis and judgment.

[0037] It should be noted that there seems to be an incomplete sentence in the translation of . The Chinese part "线下实体门店可以通过在线服务平台01完成门店信息注册、展示及更新,例如,在线服务平台01可以是地图类应用 / 平台、外卖类应用 / 平台等。" is not fully translated in the provided English text. It may need to be corrected and completed for a more accurate translation. The corrected and completed translation of should be: As shown, in fields such as local life services and retail chain operations that rely on store data support, there is a three-tier architecture relationship including an online service platform 01, an operation platform 02, and a decision-making end 03. Among them, the online service platform 01 refers to an open service platform facing users and merchants. Offline physical stores can complete store information registration, display, and update through the online service platform 01. For example, the online service platform 01 can be a map application / platform, a food delivery application / platform, etc. The operation platform 02 is a backend operation support platform for managing the online service platform 01. Its core responsibility is to integrate the store data on the online service platform and provide data support for the decision-making end 03. The three form a complete link where store information is uploaded to the online service platform, the operation platform synchronously integrates the data, and the decision-making end conducts operation decisions based on the data.In real-world scenarios, stores can independently update their information through the online service platform 01 based on operational needs (such as relocation, brand upgrades, and adjustments to business hours). However, due to the lack of an efficient and automated mechanism for data synchronization between the operations platform 02 and the online service platform 01, information asynchrony issues frequently occur. For example, a store may have already updated its information on the online service platform 01, but the corresponding original data stored in the offline database of the operations platform 02 may not be synchronized in a timely manner, resulting in a discrepancy between the information in the offline database and the latest information displayed on the online service platform 01. This discrepancy directly transmits to the decision-making end 03, triggering a series of chain reactions. For instance, routes planned by the decision-making end 03 based on outdated addresses may become inaccurate, marketing strategies based on outdated addresses may fail to reach customers, and regional service coverage analysis results based on distorted data may deviate from reality. This not only reduces user trust in the online service platform but also wastes operational resources.

[0038] To address the aforementioned issues, a data preservation process for offline databases is proposed. This process involves periodically and specifically verifying, updating, supplementing, and correcting store data stored in the offline database of the operations platform to maintain its accuracy and ensure synchronization between the store data and the latest store status on the online service platform. Additionally, data preservation can also include handling duplicate store data in the offline database.

[0039] In related technologies, data preservation primarily relies on manual annotation, which involves manually verifying, entering, and updating store data in offline databases. However, with the expansion of online service platforms and the explosive growth in the number of participating stores, the frequency of store information updates has also increased significantly. The manual annotation method is inefficient and struggles to meet the batch processing needs of large-scale store data, resulting in long data preservation cycles. The lag in manual updates further exacerbates the information asynchrony between offline databases and online service platforms. This manual annotation method has become a bottleneck restricting the improvement of operational platform management efficiency and the accuracy of decision-making.

[0040] Therefore, there is an urgent need for a more efficient data preservation and processing solution to meet the processing needs of large-scale store data, reduce reliance on manual labor, and provide reliable data support for decision-making.

[0041] To address the issues of delayed updates and high reliance on manual intervention in offline store data updates, this specification provides a store data processing method and apparatus through one or more embodiments. The method first uses a semantic similarity algorithm to identify store data from an online service platform or offline database that belongs to the same entity as the target store in the offline database. For fuzzy matching scenarios where semantic similarity fails to identify the store, multiple artificial intelligences can be introduced for collaborative judgment. Leveraging the deep analysis capabilities of AI on multi-dimensional store data, the risk of misjudgment by a single identification method is avoided, accurately solving the problem of store identity identification in complex scenarios. Finally, after confirming that the stores belong to the same entity through semantic similarity verification or AI collaborative judgment, the corresponding store data in the offline database can be processed based on the identified store data, achieving automated and accurate updates of store data, reducing reliance on manual intervention, and improving data timeliness and accuracy.

[0042] The following describes embodiments of the store data processing method provided in one or more embodiments of this specification.

[0043] See Figure 2 , Figure 2 This is a flowchart illustrating a store data processing method provided in one or more embodiments of this specification. The method can be applied to servers, terminal devices, such as… Figure 1 The example shown is an operating platform, etc. The following explanation uses an application to an operating platform as an example to illustrate this method. Figure 1 As shown, the method may include the following steps: Step S102: Obtain the first store data of the target store from the offline database.

[0044] In one or more embodiments of this specification, the offline database of the operating platform is used to store store data, providing data support for various decision-making processes. To improve the accuracy and timeliness of the store data in the offline database, data preservation processing can be performed on some or all of the store data in the offline database according to a preset time period.

[0045] In one possible implementation, the operations platform can combine the decision-making needs of the decision-makers and select store data corresponding to the stores to be processed from the offline database for data preservation. For example, the operations platform can provide a variety of data tag options for decision-makers to choose from based on their current decision-making needs. After logging into the operations platform, the platform can visualize the selectable data tags, and in response to the user's selection, the platform will trigger a store screening process.

[0046] For example, the operating platform can first receive the selection operation of data tags, where different data tags correspond to different store filtering conditions; then, the operating platform filters stores that meet the conditions from the offline database according to the store filtering conditions corresponding to the data tags selected by the user, forming a set of stores to be processed; subsequently, it retrieves the store data corresponding to each store in the set of stores to be processed from the offline database and carries out subsequent data preservation processing.

[0047] Optionally, data tags may include inspection results, recent freshness time, recent exposure data, number of transactions per customer, etc. Among them, inspection results indicate whether data has undergone freshness processing within a preset time period.

[0048] Optionally, store filtering criteria may include at least one of the following: (1) Stores whose recent freshness time is longer than the preset time; (2) Stores that have not been verified by the online service platform; (3) It belongs to a popular store; It should be noted that online service platforms typically require verification when registering a store to confirm information such as the store name and address. In some special cases, stores that have not undergone verification by the online service platform may have their data's authenticity questionable and should be given priority for data preservation.

[0049] For example, if the data tag selected by the decision-making user is the most recent shelf life, the operation platform will filter out stores with a most recent shelf life of more than 30 days to form a set of stores to be processed.

[0050] For another example, if the data tags selected by the decision-making user are inspection results, recent exposure data, and number of transactions per customer, the operation platform can determine whether each store is a hot store based on the recent exposure data and the number of transactions per customer. Then, it can filter out the hot stores and the stores whose recent freshness time is greater than the preset time to form a set of stores to be processed.

[0051] Thus, given the diversity of decision-making scenarios, the corresponding store data needs also vary. This approach offers multi-dimensional data tag options, allowing decision-makers to freely combine and select tags according to different scenarios, enabling flexible customization of filtering rules. This method uses the actual needs of the decision-makers as the core anchor, filtering stores highly relevant to the current decision through data tags. For example, when decision-makers need to conduct marketing campaigns in hot areas, they can select the "recent exposure data + number of transactions per customer" tag to accurately filter out hot stores for data preservation, ensuring that data such as store addresses, upon which marketing decisions rely, are up-to-date. To address the issue of data distortion due to long-term lack of updates, the "recent preservation time" tag can be selected to specifically process stores that have not been updated for more than 30 days. This store data selection method, guided by decision-making needs, directly links data preservation processing to the decision-making scenario, avoiding the omission of core data or the waste of resources due to indiscriminate preservation, significantly improving the accuracy of data-driven judgments by the decision-makers.

[0052] In one possible implementation, the operating platform can also determine the set of stores to be processed from an offline database based on preset store filtering rules, without requiring decision-makers to manually select tags.

[0053] For example, the preset store filtering rules can be set to automatically include stores that meet one or more of the following three conditions: the most recent freshness time exceeds a preset duration, they have not been reviewed by the online service platform, and they are considered "hot" stores. In this way, the operations platform can achieve batch and automated filtering of stores to be processed, further improving the efficiency of data freshness processing and adapting to the management needs of large-scale store data.

[0054] In this way, through the two flexible screening methods mentioned above, the operating platform can selectively choose high-priority, high-demand stores for data preservation processing, which not only ensures the quality of core decision-making data, but also avoids the waste of resources caused by indiscriminate processing, thus achieving precision and efficiency in data preservation.

[0055] It should be noted that after obtaining the set of stores to be processed, data preservation processing can be performed on the store data of each store in the set. For ease of description, the following illustrative example only illustrates the data preservation process for one store in the set. In one or more embodiments of this specification, the store is referred to as the target store, and the store data corresponding to the target store is referred to as the first store data.

[0056] The first store data can include different information depending on the data preservation target. For example, the first store data may include the store name, store address, and store latitude and longitude. Of course, the first store data may also include more or better store information, and this specification does not limit this in one or more embodiments.

[0057] Step S104: Obtain the set of similar stores corresponding to the target store.

[0058] As can be seen from the above description, data preservation processing can include two aspects: data update processing and duplicate data processing.

[0059] In the data update processing scenario, a set of similar stores corresponding to the target store can be obtained from the online service platform. Then, store data belonging to the same store as the target store can be identified from the set of similar stores to update the first store data in the offline database, so that the store data of the target store in the offline database is consistent with the latest store information displayed on the online service platform.

[0060] For scenarios involving duplicate data processing, a set of similar stores corresponding to the target store can be obtained from an offline database. Then, store data belonging to the same store as the target store is identified from the set of similar stores, and duplicate store data that are found to be duplicates of the first store data is deduplicated.

[0061] The following examples illustrate how to obtain a set of similar stores corresponding to a target store for two different data preservation processing scenarios.

[0062] like Figure 3 and Figure 4 As shown, for data update scenarios, obtaining the set of similar stores corresponding to the target store from the online service platform can be achieved through the following steps: Step S201: Obtain the latitude and longitude information corresponding to each transaction data of the target store within a preset time period.

[0063] For example, taking the target store as A Milk Tea (Science and Technology Park Store), the first store data includes the store name: A Milk Tea (Science and Technology Park Store), and the address: Room 101, Building A, Science and Technology Park, High-tech Zone, City B.

[0064] The operating platform extracted 100 transaction data records from the target store over the past 30 days from the offline database. Each transaction data record included the location latitude and longitude of the user when placing the order. These latitude and longitude records surrounded the area around Building A of the Science and Technology Park.

[0065] Step S202: Cluster the latitude and longitude information.

[0066] One possible implementation is to use the DBSCAN (Density-based Noise-based Application Spatial Clustering) clustering algorithm to cluster the latitude and longitude of 100 transactions to obtain a core cluster.

[0067] For example, the latitude and longitude points of each transaction are traversed to determine whether it is a core point. For instance, among 92 concentrated points, each point has ≥20 transactions within its 100-meter neighborhood, thus it is considered a core point, and their neighborhoods overlap, forming a core cluster. The 8 scattered points have <20 transactions within their 100-meter neighborhoods, thus they are considered noise points and do not possess core point attributes. In this way, DBSCAN automatically assigns core points to a cluster, while noise points do not belong to any cluster.

[0068] One possible implementation is to use the K-Means clustering algorithm to cluster the latitude and longitude of 100 transactions to obtain a preset number of clusters.

[0069] For example, two clusters are obtained: Cluster 1 contains the latitude and longitude of 92 transactions, which are concentrated near the actual location of Building A in the Science Park, and Cluster 2 contains the latitude and longitude of 8 transactions, which are scattered within a 3-kilometer radius.

[0070] It should be noted that the above example only illustrates the clustering result by including one or more clusters, and does not imply a limitation on the clustering result. For example, the clustering result may not even yield clusters.

[0071] For example, in cases with limited transaction data within the last 30 days—say, only 5 transactions within the last 30 days, with latitude and longitude points scattered within a 5-kilometer radius—the number of transactions within a 100-meter neighborhood for each corresponding latitude and longitude point is less than 20, indicating no core point. Thus, the DBSCAN clustering result has no core cluster.

[0072] It should also be noted that one or more embodiments in this specification do not limit the specific clustering algorithm. For example, other clustering algorithms besides K-Means clustering and DBSCAN clustering can be used.

[0073] Step S203: Determine whether there is a target latitude and longitude that can represent the location of the target store in the clustering results.

[0074] If the clustering results include one or more clusters, it can be further determined whether there are target latitude and longitude coordinates in the clustering results that can characterize the location of the target store.

[0075] In one possible implementation, when using the DBSCAN clustering algorithm to cluster latitude and longitude information to obtain a core cluster, the latitude and longitude of the geometric center of the core cluster is used as the target latitude and longitude that can characterize the store location. If no core cluster is obtained, it is determined that there is no target latitude and longitude.

[0076] In one possible implementation, when one or more clusters are obtained using the K-Means clustering algorithm, a set rule can be established to determine whether a cluster is included that can be used to determine the target latitude and longitude. If a cluster that can be used to determine the target latitude and longitude is included, the geometric center latitude and longitude of that cluster is used as the target latitude and longitude that can characterize the store location. If no cluster that can be used to determine the target latitude and longitude is included, it is determined that there is no target latitude and longitude.

[0077] For example, the rule is set such that when the sample proportion threshold of a cluster is greater than or equal to 80%, the cluster is a cluster that can be used to determine the target latitude and longitude.

[0078] Taking clusters 1 and 2 in step S202 as examples, since the sample proportion of cluster 1 is 92%≥80%, the center latitude and longitude of cluster 1 are determined as the target latitude and longitude that can characterize the store location.

[0079] Step S204: If there is a target latitude and longitude that can represent the location of the target store in the clustering results, obtain the set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store.

[0080] This method, relying on target latitude and longitude coordinates obtained through transaction data clustering, is more suitable for mature stores with stable transaction data, such as chain brand stores and popular stores in commercial districts. In real-world scenarios, some store address information may have issues such as inconsistent wording (e.g., "Building A of Science Park" written as "Science Park Block A") or outdated addresses. Using target latitude and longitude coordinates as the matching indicator eliminates the reliance on text addresses. Even if the address information is incorrect, it can accurately retrieve similar stores based on real geographical locations, improving the accuracy of matching results.

[0081] The operating platform can search for similar stores to the target store from one or more online service platforms. For example, online service platforms can include map applications, food delivery applications, etc.

[0082] One possible implementation is, such as Figure 4 As shown, the operating platform can obtain similar stores that meet the criteria from the online service platform through a general API search.

[0083] API general search refers to a technical method that uses standardized API interfaces provided by online service platforms to retrieve a batch of store data sets that meet the conditions based on fuzzy search criteria with non-precise matching.

[0084] For example, if the clustering results contain target latitude and longitude that can characterize the location of the target store, the operation platform calls the interface of the online service platform and passes the search parameters to the interface. The search parameters include the surrounding preset geographical range centered on the target latitude and longitude, and the fuzzy matching rules of the store name; and receives the set of similar stores that meet the search parameters from the interface.

[0085] For example, the operating platform calls the API of a food delivery application (see...). Figure 4 (The search interface is called), and the following search parameters are passed in: Geographic range: 500 meters around the target's latitude and longitude; Name fuzzy matching: contains "A milk tea".

[0086] Based on these generalization conditions, the interface will return beverage stores with names containing "A Milk Tea" within the geographical range (e.g., A Milk Tea Technology Park Store, A Milk Tea Technology Park Branch, A Milk Tea (Technology Park Building A Store) etc.), forming a set of similar stores.

[0087] In this way, the API's general search uses the target's latitude and longitude along with the store's name as the retrieval basis. Through the dual constraints of geographical scope and name keyword matching, it can quickly filter out cross-regional and unrelated stores, significantly reducing the probability of false matches such as stores with the same name or different names at the same location. For example, when searching for "A Milk Tea," combining the target's latitude and longitude can accurately pinpoint the brand's branches within the same business district, avoiding matches with stores of the same name in other cities. This ensures the spatial and attribute relevance between the set of similar stores and the target store, providing a high-quality candidate data foundation for subsequent semantic similarity calculations and AI model judgments.

[0088] Furthermore, the API-based general search is implemented using the standardized API interface of the online service platform, supporting a single call to return a batch of store data that meets the criteria, eliminating the need for multiple point-to-point queries for individual stores. For the data processing needs of tens or hundreds of thousands of stores on the operations platform, this approach can significantly reduce the frequency of API calls and lower the resource consumption for network transmission and data processing.

[0089] In one possible implementation, the operating platform can also use web crawlers to selectively crawl publicly available store data to obtain a set of similar stores.

[0090] For example, the operating platform can configure the crawler's collection scope and filtering conditions to ensure the targeted nature of the collected data. For instance, a preset geographical range centered on the target's latitude and longitude can be used as the collection boundary, and fuzzy matching rules for store names can be used as filtering conditions to crawl corresponding store data from the online service platform.

[0091] In this way, for online service platforms with open interfaces, a set of similar stores corresponding to the target store can be obtained through interface-based general search. For online service platforms without open interfaces, a set of similar stores corresponding to the target store can be collected through web crawling, thus achieving full coverage of data across multiple platforms.

[0092] Step S205: If there is no target latitude and longitude that can represent the location of the target store in the clustering results, obtain the set of similar stores corresponding to the target store from the online service platform according to the store name and store address of the target store.

[0093] This method does not rely on transaction data clustering results and is more suitable for scenarios such as new stores, temporary stores, and community stores where there is insufficient transaction data to support latitude and longitude clustering. For example, a newly registered "XX Coffee Temporary Store" may not have historical transaction latitude and longitude data, but it can still effectively obtain similar stores in the vicinity by matching "name keywords + address range".

[0094] The specific implementation of retrieving the set of similar stores corresponding to the target store from the online service platform based on the target store's name and address can be found in the description of step S204, and will not be repeated here. For example, a pre-defined geographical range centered on the target's latitude and longitude, along with fuzzy matching rules for the store names, can be used as search parameters to retrieve the set of similar stores corresponding to the target store from the online service platform through a general interface search. Alternatively, a web crawler can be used, with a pre-defined geographical range centered on the store address as the collection boundary and fuzzy matching rules for the store names as filtering conditions, to crawl the corresponding store data from the online service platform.

[0095] In this way, for stores with stable transaction data and clear locations (such as chain brand directly operated stores), precise positioning is achieved by matching the target latitude and longitude with the store name. For stores with less transaction data and scattered locations (such as temporary stores and new stores), the system switches to matching the store name with the store address, covering the matching needs of different types of stores in the operating platform.

[0096] like Figure 5 As shown, for scenarios involving duplicate data processing, obtaining a set of similar stores corresponding to a target store from an offline database can be achieved using the following steps: S301, determine the store similarity between the target store and other stores in the offline database.

[0097] The above step S301 can also be referred to as offline similarity pair data calculation, where offline similarity pair data may include a set of store pairs and the similarity score of the store pairs. For example, Figure 5 The offline similarity pair data in the figure represents the similarity score between target store d1 and store d2.

[0098] For example, a Cartesian product (i.e., generating unordered pairings of any two stores) can be constructed first for all stores in the offline database. Then, for each store pair, a similarity score is calculated based on a preset similarity algorithm (such as edit distance, cosine similarity, etc.), ultimately forming a similarity matrix between all stores. Afterward, store pairs including the target store and their corresponding similarity scores can be obtained from the similarity matrix.

[0099] S302: Obtain from the offline database stores that belong to the same service scenario as the target store and whose store similarity is greater than the second similarity threshold, and obtain the set of similar stores corresponding to the target store.

[0100] After obtaining the store pairs of the target store and the corresponding similarity scores for each store pair, store pairs that meet the following criteria are selected: the other store in each store pair belongs to the same service scenario as the target store, and the similarity score is greater than the second similarity threshold. Then, the other stores in each store pair that meet the above conditions, excluding the target store, are selected as the set of similar stores.

[0101] It should be noted that the online service platform may or may not retrieve similar stores corresponding to the target store. If similar stores corresponding to the target store are retrieved, meaning the set of similar stores includes at least one similar store, the following step S106 can be executed. If no similar stores corresponding to the target store are retrieved, meaning the set of similar stores does not include any similar stores, information can be sent to the manual annotation end to instruct that the target store be manually annotated for data preservation.

[0102] If the set of similar stores includes at least one similar store, the store data in the set of similar stores can be iterated through, and steps S106 to S108 can be executed until a similar store that is the same as the target store is found. If no similar store that is the same as the target store is found after the iteration ends, information can be sent to the manual annotation terminal to instruct that the target store be manually annotated for data preservation.

[0103] Step S106: If the set of similar stores includes at least one similar store, determine the semantic similarity between the target store and the target similar store based on the first store data and the second store data.

[0104] The second set of store data consists of the store data corresponding to the target similar store in the similar store set on the online service platform or in the offline database. The target similar store is the currently visited similar store.

[0105] The first and second store data can include store data from multiple dimensions, which may be data requiring preservation. For example, both the first and second store data may include the store's name, address, and latitude and longitude.

[0106] In one possible implementation, the semantic similarity between the target store and the target similar stores is determined based on the data of the first store and the data of the second store. This can be achieved as follows: when the set of similar stores includes at least one similar store, the single-dimensional similarity between the data of the first store and the data of the second store is calculated for each dimension of the store data. Then, based on the weight values ​​corresponding to each dimension, the single-dimensional similarity corresponding to each dimension is weighted to obtain the semantic similarity between the target store and the target similar stores.

[0107] For example, the similarity of store names, addresses, and latitude / longitude is calculated between the first store data and the second store data to obtain the single-dimensional similarity for each dimension. Then, based on the weight values ​​corresponding to the store names, addresses, and latitude / longitude, the single-dimensional similarity corresponding to the store names, addresses, and latitude / longitude is weighted to obtain the semantic similarity between the target store and the target similar stores.

[0108] It should be noted that the specific implementation method for calculating single-dimensional similarity is not limited in one or more embodiments of this specification. For example, single-dimensional similarity can be calculated using methods such as edit distance method and cosine similarity method.

[0109] In one possible implementation, the semantic similarity between the target store and the target similar stores is determined based on the data of the first store and the data of the second store. This can be achieved by inputting the data of the first store and the data of the second store into a pre-trained semantic similarity model, and then outputting the semantic similarity between the target store and the target similar stores through the semantic similarity model.

[0110] If the semantic similarity is greater than or equal to the first similarity threshold, it can be determined that the target store and the target similar store are the same store, and step S110 is executed. If the semantic similarity is less than the first similarity threshold, step S108 can be executed to further determine whether the target store and the target similar store are the same store.

[0111] In one possible implementation, if the weight value corresponding to latitude and longitude in the semantic similarity calculation is less than the weight threshold, the following steps may be included: if the semantic similarity is greater than or equal to the first similarity threshold, determine whether the distance between the target store and the target similar store is less than or equal to a distance threshold based on the latitude and longitude of the target store and the target similar store; if the distance is less than or equal to the distance threshold, it can be determined that the target store and the target similar store are the same store, and then step S110 is executed. If the distance is greater than the distance threshold, step S108 may be executed to further determine whether the target store and the target similar store are the same store.

[0112] It should be noted that, when a target latitude and longitude exist, the distance between the target store and similar stores is determined based on the target latitude and longitude and the latitude and longitude in the second store data, to see if the distance is less than or equal to a distance threshold. When a target latitude and longitude do not exist, the distance between the target store and similar stores is determined based on the latitude and longitude in the first store data and the latitude and longitude in the second store data, to see if the distance is less than or equal to a distance threshold.

[0113] Step S108: If the semantic similarity is less than the first similarity threshold, use multiple artificial intelligence models to identify whether the first store data and the second store data belong to the same store.

[0114] To prevent misclassification of similar stores belonging to the same store as different stores based on semantic similarity, multiple artificial intelligence models are further used to identify whether the similar store and the target store are the same store when the semantic similarity is less than the first similarity threshold.

[0115] For example, first store data, second store data, and prompt words can be input into multiple different large language models. The prompt words are used to indicate whether the target similar store and the target store are the same store in the large language pattern recognition. For example, the prompt words may include task instructions, rule constraints, data samples, output format requirements, etc., which are not limited in this specification in one or more embodiments.

[0116] Afterwards, each language mode outputs its own recognition result based on the above input. The recognition result may include: the same store or different stores.

[0117] Finally, based on the recognition results of various language models and preset voting rules, it can be determined whether the target store and similar stores are the same store.

[0118] In one possible implementation, determining whether a target store and a similar store are the same store, based on the recognition results of various language models and preset voting rules, can be achieved as follows: Count the number of target recognition results among the recognition results of various language models, where each target recognition result represents a store that is the same as a similar store. Then, calculate the percentage of the target recognition results out of the total number of recognition results. If the percentage exceeds a preset threshold, the target store and the similar store are determined to be the same store. If the percentage does not exceed the preset threshold, the target store and the similar store are determined to be different stores.

[0119] One possible implementation involves determining whether a target store and a similar store are the same store based on the recognition results of various language models and preset voting rules. This can be achieved as follows: Based on historical validation data of each language model (such as past accuracy and recall), assign weight values ​​to each participating language model in determining store identity. Quantize the recognition results output by each language model; for example, assign a quantized value of 1 for a result indicating the same store and a quantized value of 0 for a result indicating different stores. Then, calculate a comprehensive score based on the weight values ​​and quantized values ​​of each model. If the comprehensive score is greater than a preset threshold, determine that the target store and the similar store are the same store; otherwise, determine that they are not the same store. The weight values ​​of each language model can then be dynamically optimized based on this determination result.

[0120] In this way, the method combines the characteristics and historical performance of different large language models, assigns differentiated weights to each large language model, and continuously adapts to changes in the capabilities of large language models through weight iteration, thus ensuring the long-term stability of the judgment results.

[0121] Step S110: Using multiple artificial intelligence models, if the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to the first similarity threshold, the first store data in the offline database is processed based on the second store data.

[0122] If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to the first similarity threshold, then the target similar store traversed this time is considered to be the same store as the target store, and the traversal of other similar stores in the similar store set can be ended.

[0123] For data update processing scenarios, it is recommended to use the second store's data to update the first store's data in the offline database.

[0124] For example, the store name, store address, and store latitude and longitude coordinates from the second store data can be used as the recommended name, recommended address, and recommended latitude and longitude coordinates for subsequent updates to the first store data in the offline database. For instance, the store name, store address, and store latitude and longitude coordinates from the first store data in the offline database can be updated to the recommended name, recommended address, and recommended latitude and longitude coordinates.

[0125] For scenarios involving duplicate data processing, duplicate data from the first store and the second store in the offline database can be deduplicated.

[0126] Thus, the data processing method provided in one or more embodiments of this specification can first identify store data belonging to the same entity as the target store in the offline database from an online service platform or offline database using a semantic similarity algorithm. For fuzzy matching scenarios where semantic similarity fails to identify the store, multiple artificial intelligence models can be introduced for collaborative judgment. Leveraging the deep analysis capabilities of these models on multi-dimensional store data, the risk of misjudgment by a single identification method can be avoided, accurately solving the problem of store identity identification in complex scenarios. Finally, after confirming that the stores belong to the same entity through semantic similarity verification or collaborative judgment by multiple artificial intelligence models, the corresponding store data in the offline database can be processed based on the identified store data, achieving automated and accurate updates of store data, improving data timeliness and accuracy, and reducing reliance on manual intervention.

[0127] It is understood that the above embodiments are merely examples, and modifications can be made to the above embodiments in actual implementation. Those skilled in the art will understand that any modifications to the above embodiments that do not require creative effort fall within the protection scope of one or more embodiments of this specification, and will not be described again in the embodiments.

[0128] Based on the same inventive concept, one or more embodiments of this specification also provide a store data processing device. Since the principle of the problem solved by the store data processing device is similar to that of the aforementioned store data processing method, the implementation of the store data processing device can refer to the implementation of the aforementioned store data processing method, and the repeated parts will not be described again.

[0129] See Figure 6 , Figure 6 This is a structural block diagram of a store data processing device provided for one or more embodiments of this specification. Figure 6 As shown, the store data processing device 400 may include: a first acquisition module 401, a second acquisition module 402, a determination module 403, an identification module 404, and a data processing module 405. Among them, The first acquisition module 401 is used to acquire the first store data of the target store from the offline database; the second acquisition module 402 is used to acquire the set of similar stores corresponding to the target store. The determining module 403 is used to determine the semantic similarity between the target store and the target similar store based on the first store data and the second store data when the similar store set includes at least one similar store, wherein the second store data is the store data corresponding to the target similar store in the similar store set. The identification module 404 is used to identify whether the first store data and the second store data belong to the same store by using multiple artificial intelligence models when the semantic similarity is less than the first similarity threshold. The data processing module 405 is used to process the first store data in the offline database based on the second store data when multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or when the semantic similarity is greater than or equal to a first similarity threshold.

[0130] In one possible implementation, the second acquisition module 402 is configured to: acquire the latitude and longitude information corresponding to each transaction data of the target store within a preset time period; perform clustering processing on the latitude and longitude information; if there is a target latitude and longitude that can characterize the location of the target store in the clustering results, acquire a set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store; if there is no target latitude and longitude that can characterize the location of the target store in the clustering results, acquire a set of similar stores corresponding to the target store from the online service platform according to the store name and store address of the target store; wherein, the second store data is the store data of the target similar stores in the similar store set on the online service platform.

[0131] In one possible implementation, the second acquisition module 402 is used to, when the clustering results contain target latitude and longitude that can characterize the location of the target store, obtain a set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store. Specifically, when the clustering results contain target latitude and longitude that can characterize the location of the target store, the module calls the interface of the online service platform and passes search parameters to the interface. The search parameters include a preset geographical range centered on the target latitude and longitude and a fuzzy matching rule for the store name; and receives the set of similar stores that satisfy the search parameters from the interface.

[0132] In one possible implementation, the data processing module 405 is configured to: update the first store data in the offline database based on the second store data, provided that the first store data and the second store data belong to the same store, or the semantic similarity is greater than or equal to a first similarity threshold, using multiple artificial intelligence models.

[0133] In one possible implementation, the second acquisition module 402 is used to: determine the store similarity between the target store and other stores in the offline database; acquire stores from the offline database that belong to the same service scenario as the target store and have a store similarity greater than a second similarity threshold, thereby obtaining a set of similar stores corresponding to the target store; wherein, the second store data is the store data of the target similar store in the offline database.

[0134] In one possible implementation, the data processing module 405 is used to deduplicate the first store data and the second store data in the offline database when multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or when the semantic similarity is greater than or equal to a first similarity threshold.

[0135] In one possible implementation, the first acquisition module 401 is configured to: receive a selection operation for data tags, wherein different data tags correspond to different store filtering conditions; determine a set of stores to be processed based on the store filtering conditions corresponding to the selected data tags, wherein the set of stores to be processed includes the target store; and acquire the first store data of the target store from an offline database.

[0136] In one possible implementation, the first store data and the second store data include store data in multiple dimensions; wherein, the determining module 403 is configured to: when the set of similar stores includes at least one similar store, calculate the single-dimensional similarity between the first store data and the second store data in the corresponding dimension for each dimension of the store data; and perform weighted processing on the single-dimensional similarity corresponding to each dimension based on the weight values ​​corresponding to each dimension to obtain the semantic similarity between the target store and the target similar store.

[0137] In one possible implementation, the recognition module 404 is used to: count the number of target recognition results in the recognition results of various language models, wherein the target recognition result is the recognition result that indicates that the target store and the target similar store are the same store; calculate the proportion of the number of target recognition results to the total number of recognition results; and determine that the target store and the target similar store are the same store if the proportion exceeds a preset proportion threshold.

[0138] See Figure 7 , Figure 7 This is a structural block diagram of an electronic device provided for one or more embodiments of this specification. Figure 7 As shown, the electronic device 500 may include a processor 501 and a memory 502; the memory 502 may be coupled to the processor 501. It is worth noting that... Figure 7 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.

[0139] In one possible implementation, the functionality of the store data processing device 400 can be integrated into the processor 501. The processor 501 can be configured to perform the following operations: Retrieve the first store data of the target store from the offline database; Obtain the set of similar stores corresponding to the target store; When the set of similar stores includes at least one similar store, the semantic similarity between the target store and the target similar store is determined based on the first store data and the second store data, wherein the second store data is the store data corresponding to the target similar store in the set of similar stores; If the semantic similarity is less than the first similarity threshold, multiple artificial intelligence models are used to identify whether the first store data and the second store data belong to the same store. If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then the first store data in the offline database is processed based on the second store data.

[0140] In another possible implementation, the store data processing device 400 can be configured separately from the processor 501. For example, the store data processing device 400 can be configured as a chip connected to the processor 501, and store data processing can be achieved through the control of the processor 501.

[0141] Furthermore, in some alternative implementations, the electronic device 500 may also include: a communication module, an input unit, an audio processor, a display, a power supply, etc. It is worth noting that the electronic device 500 is not necessarily required to include these components. Figure 7 All components shown; in addition, the electronic device 500 may also include Figure 7 For components not shown, please refer to existing technologies.

[0142] In some alternative implementations, the processor 501, sometimes also referred to as a controller or operating control, may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of the electronic device 500.

[0143] The memory 502 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable devices. It may store the aforementioned information related to the store data processing device 400, and may also store programs for executing that information. The processor 501 may execute the program stored in the memory 502 to perform information storage or processing, etc.

[0144] An input unit can provide input to the processor 501. This input unit may be, for example, a button or touch input device. A power supply can be used to provide power to the electronic device 500. A display can be used to display images and text, etc. This display may be, for example, an LCD display, but is not limited to this.

[0145] Memory 502 can be a solid-state memory, such as read-only memory (ROM), random access memory (RAM), SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROM, etc. Memory 502 can also be some other type of device. Memory 502 includes buffer memory (sometimes referred to as a buffer). Memory 502 may include an application / function storage unit for storing application programs and function programs or processes for executing the operation of electronic device 500 via processor 501.

[0146] The memory 502 may also include a data storage unit for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit of the memory 502 may include various drivers for the computer device for communication functions and / or for performing other functions of the computer device (such as messaging applications, address book applications, etc.).

[0147] The communication module is a transmitter / receiver that sends and receives signals via an antenna. The communication module (transmitter / receiver) is coupled to the processor 501 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.

[0148] Based on different communication technologies, multiple communication modules can be configured in the same computer device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication modules (transmitters / receivers) are also coupled to speakers and microphones via an audio processor to provide audio output through the speakers and receive audio input from the microphones, thereby enabling typical telecommunications functions. The audio processor may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processor is coupled to processor 501, enabling on-device recording via the microphone and on-device playback of stored sound via the speakers.

[0149] One or more embodiments of this specification also provide a computer-readable storage medium capable of implementing all steps of the store data processing method in the above embodiments. The computer-readable storage medium stores a computer program that, when executed by a processor, implements all steps of the store data processing method in the above embodiments. For example, when the processor executes the computer program, it implements the following steps: Retrieve the first store data of the target store from the offline database; Obtain the set of similar stores corresponding to the target store; When the set of similar stores includes at least one similar store, the semantic similarity between the target store and the target similar store is determined based on the first store data and the second store data, wherein the second store data is the store data corresponding to the target similar store in the set of similar stores; If the semantic similarity is less than the first similarity threshold, multiple artificial intelligence models are used to identify whether the first store data and the second store data belong to the same store. If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then the first store data in the offline database is processed based on the second store data.

[0150] While one or more embodiments of this specification provide the operational steps of the methods described in the embodiments or flowcharts, more or fewer operational steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual device or client product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).

[0151] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, apparatus (systems), or computer program products. Therefore, the embodiments of this specification can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, one or more embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] This specification describes one or more embodiments of a method, apparatus (system), and computer program product according to one or more embodiments of this specification with reference to flowchart illustrations and / or block diagrams. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0153] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 Figure 1 The steps of the function specified in one or more boxes.

[0155] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device and system embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0156] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Those skilled in the art will understand the specific meaning of the above terms in one or more embodiments of this specification, depending on the specific circumstances.

[0157] It should be noted that, unless otherwise specified, one or more embodiments and features thereof in this specification can be combined with each other. This specification is not limited to any single aspect, nor to any single embodiment, nor to any combination and / or substitution of such aspects and / or embodiments. Furthermore, each aspect and / or embodiment of one or more embodiments of this specification can be used alone or in combination with one or more other aspects and / or embodiments thereof.

[0158] It should also be noted that the user / store information and data involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties.

[0159] The data acquisition and collection actions involved in one or more embodiments of this specification are all performed after authorization by the user, the object, or after full authorization by all parties.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of one or more embodiments of this specification, and are not intended to limit them. Although one or more embodiments of this specification have been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of one or more embodiments of this specification, and they should all be covered within the scope of the claims and the specification of one or more embodiments of this specification.

[0161] The foregoing description of one or more embodiments of this specification has been provided in conjunction with optional implementation methods. However, these embodiments are merely exemplary and serve only an illustrative purpose. Based on this, various substitutions and modifications can be made to one or more embodiments of this specification, all of which fall within the protection scope of one or more embodiments of this specification.

Claims

1. A method for processing store data, characterized in that, The method includes: Retrieve the first store data of the target store from the offline database; Obtain the set of similar stores corresponding to the target store; When the set of similar stores includes at least one similar store, the semantic similarity between the target store and the target similar store is determined based on the first store data and the second store data, wherein the second store data is the store data corresponding to the target similar store in the set of similar stores; If the semantic similarity is less than the first similarity threshold, multiple artificial intelligence models are used to identify whether the first store data and the second store data belong to the same store. If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then the first store data in the offline database is processed based on the second store data.

2. The method as described in claim 1, characterized in that, Obtain the set of similar stores corresponding to the target store, including: Obtain the latitude and longitude information corresponding to each transaction data of the target store within a preset time period; The latitude and longitude information is then clustered. If the clustering results contain a target latitude and longitude that can characterize the location of the target store, then obtain the set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store; If no target latitude and longitude coordinates that can characterize the location of the target store are found in the clustering results, the set of similar stores corresponding to the target store is obtained from the online service platform according to the store name and store address of the target store. The second store data refers to the store data of the target similar stores in the similar store set on the online service platform.

3. The method as described in claim 2, characterized in that, If the clustering results contain target latitude and longitude coordinates that can characterize the location of the target store, then, based on the target latitude and longitude coordinates and the store name of the target store, obtain a set of similar stores corresponding to the target store from the online service platform, including: If the clustering results contain a target latitude and longitude that can characterize the location of the target store, call the interface of the online service platform and pass the search parameters to the interface. The search parameters include a preset geographical range centered on the target latitude and longitude, and a fuzzy matching rule for the store name. Receive the set of similar stores that satisfy the search parameters from the interface.

4. The method as described in claim 2, characterized in that, If, using multiple artificial intelligence models, the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then data processing is performed on the first store data in the offline database based on the second store data, including: If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, the first store data in the offline database is updated based on the second store data.

5. The method as described in claim 1, characterized in that, Obtain the set of similar stores corresponding to the target store, including: Determine the store similarity between the target store and other stores in the offline database; Obtain stores from the offline database that belong to the same service scenario as the target store and whose store similarity is greater than the second similarity threshold, and obtain a set of similar stores corresponding to the target store; The second store data refers to the store data of the target similar stores in the similar store set in the offline database.

6. The method as described in claim 5, characterized in that, If, using multiple artificial intelligence models, the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then data processing is performed on the first store data in the offline database based on the second store data, including: If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, then the first store data and the second store data in the offline database are deduplicated.

7. The method according to any one of claims 1 to 6, characterized in that, Retrieve the first store data of the target store from the offline database, including: Receive the selection operation of data tags, where different data tags correspond to different store filtering conditions; Based on the store filtering conditions corresponding to the selected data tags, a set of stores to be processed is determined, including the target stores. Obtain the first store data of the target store from the offline database.

8. The method according to any one of claims 1 to 6, characterized in that, The first store data and the second store data include store data from multiple dimensions; Where the set of similar stores includes at least one similar store, determining the semantic similarity between the target store and the target similar store based on the first store data and the second store data includes: When the set of similar stores includes at least one similar store, the single-dimensional similarity between the first store data and the second store data in the corresponding dimension is calculated for each dimension of the store data. Based on the weight values ​​corresponding to each dimension, the single-dimensional similarity corresponding to each dimension is weighted to obtain the semantic similarity between the target store and the target similar store.

9. The method according to any one of claims 1 to 6, characterized in that, Using multiple artificial intelligence models, the system identifies whether the first store data and the second store data belong to the same store, including: The first store data, the second store data, and prompt words are respectively input into multiple large language models. The prompt words are used to instruct the large language models to identify whether the two stores are the same store. Based on the recognition results of various language models and preset voting rules, it is determined whether the target store and the target similar store are the same store.

10. A store data processing device, characterized in that, The device includes: The first acquisition module is used to acquire the first store data of the target store from the offline database; The second acquisition module is used to acquire a set of similar stores corresponding to the target store; The determining module is used to determine the semantic similarity between the target store and the target similar store based on the first store data and the second store data when the similar store set includes at least one similar store, wherein the second store data is the store data corresponding to the target similar store in the similar store set; The identification module is used to identify whether the first store data and the second store data belong to the same store by using multiple artificial intelligence models when the semantic similarity is less than a first similarity threshold. The data processing module is used to process the first store data in the offline database based on the second store data, when multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or when the semantic similarity is greater than or equal to a first similarity threshold.

11. The apparatus as claimed in claim 10, characterized in that, The second acquisition module is used for: Obtain the latitude and longitude information corresponding to each transaction data of the target store within a preset time period; The latitude and longitude information is clustered. If the clustering results contain a target latitude and longitude that can characterize the location of the target store, then obtain the set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store; If no target latitude and longitude coordinates that can characterize the location of the target store are found in the clustering results, the set of similar stores corresponding to the target store is obtained from the online service platform according to the store name and store address of the target store. The second store data refers to the store data of the target similar stores in the similar store set on the online service platform.

12. The apparatus as claimed in claim 11, characterized in that, The second acquisition module is used to, when the clustering results contain target latitude and longitude that can characterize the location of the target store, obtain a set of similar stores corresponding to the target store from the online service platform according to the target latitude and longitude and the store name of the target store, specifically: If the clustering results contain a target latitude and longitude that can characterize the location of the target store, call the interface of the online service platform and pass the search parameters to the interface. The search parameters include a preset geographical range centered on the target latitude and longitude, and a fuzzy matching rule for the store name. Receive the set of similar stores that satisfy the search parameters from the interface.

13. The apparatus as claimed in claim 11, characterized in that, The data processing module is used for: If multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or if the semantic similarity is greater than or equal to a first similarity threshold, the first store data in the offline database is updated based on the second store data.

14. The apparatus as claimed in claim 10, characterized in that, The second acquisition module is used for: Determine the store similarity between the target store and other stores in the offline database; Obtain stores from the offline database that belong to the same service scenario as the target store and whose store similarity is greater than the second similarity threshold, and obtain a set of similar stores corresponding to the target store; The second store data refers to the store data of the target similar stores in the similar store set in the offline database.

15. The apparatus as claimed in claim 14, characterized in that, The data processing module is used to deduplicate the first store data and the second store data in the offline database when multiple artificial intelligence models are used to identify that the first store data and the second store data belong to the same store, or when the semantic similarity is greater than or equal to a first similarity threshold.

16. The apparatus as claimed in any one of claims 10 to 15, characterized in that, The first acquisition module is used for: Receive the selection operation of data tags, where different data tags correspond to different store filtering conditions; Based on the store filtering conditions corresponding to the selected data tags, a set of stores to be processed is determined, including the target stores. Obtain the first store data of the target store from the offline database.

17. The apparatus as claimed in any one of claims 10 to 15, characterized in that, The first store data and the second store data include store data across multiple dimensions; wherein, the determining module is used for: When the set of similar stores includes at least one similar store, the single-dimensional similarity between the first store data and the second store data in the corresponding dimension is calculated for each dimension of the store data. Based on the weight values ​​corresponding to each dimension, the single-dimensional similarity corresponding to each dimension is weighted to obtain the semantic similarity between the target store and the target similar store.

18. The apparatus as claimed in any one of claims 10 to 15, characterized in that, The recognition module is used for: The first store data, the second store data, and prompt words are respectively input into multiple large language models. The prompt words are used to instruct the large language models to identify whether the two stores are the same store. Based on the recognition results of various language models and preset voting rules, it is determined whether the target store and the target similar store are the same store.

19. An electronic device, characterized in that, The electronic device includes: Memory, used to store computer program products; A processor is configured to execute a computer program product stored in the memory, wherein, when the computer program product is executed, it implements the method described in any one of claims 1-9.

20. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed, implement the method described in any one of claims 1-9.