A method for normalizing multi-platform stores
By adopting a hybrid cleaning strategy based on rules engine and deep learning model and multimodal feature extraction technology in a multi-platform environment, the store information is intelligently standardized and efficiently matched, which solves the problem that store information is difficult to normalize and match in a multi-platform environment, and achieves efficient and accurate data matching and data security.
Patent Information
- Application Number
- CN202510295005.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-13
AI Technical Summary
In a multi-platform environment, it is difficult to intelligently normalize and cannot be matched efficiently, resulting in difficulty in identifying, data redundancy and inefficient operation.
The original data of stores from multiple platforms in real time is captured through the API interface, and a hybrid cleaning strategy based on the rules engine and deep learning model is adopted to intelligently standardize the data. Then, a multi-dimensional feature vector set is generated using multi-modal feature extraction technology, and parallel matching calculation is performed through a multi-algorithm fusion engine, a dynamic weight allocation mechanism is designed, and a normalized platform ID and cross-platform mapping relationship table are output. Finally, the normalized results are stored in a distributed database based on blockchain for incremental synchronization and data storage.
It realizes intelligent standardized processing and efficient matching of original data in multi-platform stores, solves the problems of identification difficulties, data redundancy and low operational efficiency, significantly improves the accuracy of cross-platform data matching, and enhances the security and immutability of data through blockchain technology.
Smart Images

Figure CN119807399B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for normalizing multi-platform stores. Background Art
[0002] In today's multi-platform business scenarios such as e-commerce, food delivery, and offline retail, the same store may use different names, logos, or description information on different platforms (such as food delivery platforms, navigation platforms, social media platforms, etc.). Due to differences in the rules, naming habits, and data sources of each platform, companies face problems such as difficulty in identification, data redundancy, and low operational efficiency when managing store information on multiple platforms. Existing technologies often rely on manual comparison or simple automated scripts for processing, which is difficult to meet the needs of rapidly changing data environments. Especially when faced with massive amounts of heterogeneous data, traditional methods are not only inefficient, but also difficult to ensure data consistency and accuracy. In addition, most of the existing data cleaning and matching technologies are based on fixed rules and lack understanding and support for the semantic level, especially when processing multilingual and unstructured data, which further limits their scope of application and effectiveness. Summary of the invention
[0003] In view of the above existing problems, the present invention is proposed.
[0004] Therefore, the present invention provides a multi-platform store normalization method to solve the problem of difficulty in intelligent normalization and inefficient matching of store information in a multi-platform environment.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0006] In a first aspect, the present invention provides a method for unifying stores on multiple platforms, which comprises:
[0007] The original store data of multiple platforms is captured in real time through the API interface, and a hybrid cleaning strategy based on the rule engine and the deep learning model is adopted to perform intelligent standardization processing on the original store data of multiple platforms to obtain a standardized data set; the multiple platforms include food delivery platforms, navigation platforms and social media;
[0008] The standardized data set is extracted through multimodal feature extraction technology to generate a multidimensional feature vector set;
[0009] Input the multi-dimensional feature vector set into the multi-algorithm fusion engine, perform parallel matching calculations, design a dynamic weight allocation mechanism, output the normalized platform ID and cross-platform mapping relationship table, and generate a matching confidence report;
[0010] The normalized results are stored in a distributed database based on blockchain, an incremental synchronization mechanism is set up, the data is stored by geographical region, and the node load is dynamically distributed through the consistent hashing algorithm to obtain the updated normalized platform data.
[0011] As a preferred solution of the method for normalizing stores on multiple platforms of the present invention, the method of real-time capturing the original store data on multiple platforms through the API interface includes the following steps:
[0012] Design a unified API interface adaptation layer for different platforms. For platforms that do not support API, use a distributed crawler system based on the Scrapy framework, configure dynamic IP proxy and request frequency control;
[0013] Capture store name, address, contact information, business scope and business hours in real time, record the data source platform, capture timestamp and data version number, and output the original data table.
[0014] As a preferred solution of the multi-platform store normalization method of the present invention, wherein: a hybrid cleaning strategy based on a rule engine and a deep learning model is adopted to perform intelligent standardization processing on the original store data of multiple platforms to obtain a standardized data set, including the following steps:
[0015] Use regular expressions to remove extra spaces and punctuation marks from the original data table, and standardize common abbreviations in the addresses to obtain a preliminarily cleaned data table;
[0016] Intelligently convert the multi-language and multi-format data in the data table after preliminary cleaning, extract store-related corpus from the public multi-language parallel corpus, and combine it with the custom industry corpus to perform sub-word segmentation, remove stop words, and unify the capitalization of the corpus;
[0017] A Seq2Seq model based on the Transformer architecture is used to input the original name and address into the encoder and output the standardized Chinese name and address through the decoder;
[0018] Input the standardized address into the geocoding service, convert it into longitude and latitude coordinates, and parse out the detailed address level information;
[0019] Integrate multi-language and multi-format data that have undergone preliminary cleaning and intelligent conversion, as well as data on address geocoding and hierarchical parsing, to form a standardized data set.
[0020] As a preferred solution of the multi-platform store normalization method of the present invention, wherein: extracting data from the standardized data set by multimodal feature extraction technology to generate a multidimensional feature vector set includes the following steps:
[0021] The pre-trained BERT multilingual model is used to extract semantic features from the name field in the standardized dataset;
[0022] Input the name field into the BERT model to obtain the context embedding vector for each character;
[0023] Use a bidirectional LSTM network to further extract contextual semantic features of the name field and generate a name semantic vector;
[0024] Extract spatial features from the latitude and longitude coordinates of the address field in the standardized dataset;
[0025] Use the Geohash algorithm to encode the longitude and latitude coordinates into a string;
[0026] Convert the Geohash string to binary code to generate a space vector;
[0027] Extract unique identifiers from the contact information field in the standardized data set, and use regular expressions to extract partial numbers of the contact information as unique identifiers;
[0028] Hash the unique identifier to generate a contact vector;
[0029] The TextCNN model based on the attention mechanism is used to extract keyword features from the business scope in the standardized data set;
[0030] The business scope text is input into the TextCNN model, which includes an embedding layer, a convolutional layer, an attention layer, and a fully connected layer;
[0031] The convolution layer uses multiple convolution kernels of different sizes to extract local features, and the attention layer calculates the importance weight of each word and outputs the keyword vector of the business scope;
[0032] The name semantic vector, space vector, contact information vector and keyword vector are integrated into a multi-dimensional feature vector set for each store, and a feature similarity calculation rule base is designed.
[0033] As a preferred solution of the multi-platform store normalization method of the present invention, the design feature similarity calculation rule includes the following steps:
[0034] Extract the semantic vectors of the names of the two stores from the multidimensional feature vector set, and use cosine similarity to measure the semantic similarity of the two store names;
[0035] Extract the Geohash codes of the two stores and measure the spatial proximity of the two store addresses by calculating the Geohash prefix overlap;
[0036] Extract the contact information vectors of the two stores and use the exact matching method to match the mobile phone numbers;
[0037] For other contact information, the Hamming distance between two contact information vectors is calculated using fuzzy hash matching;
[0038] Extract the keyword vectors of the two stores and measure the similarity of the business scope of the two stores by calculating the cosine similarity of the keyword vectors;
[0039] The name similarity, address similarity, contact information similarity and business scope similarity calculation rules are integrated into a similarity calculation rule library.
[0040] As a preferred solution of the multi-platform store normalization method of the present invention, the multi-dimensional feature vector set is input into the multi-algorithm fusion engine, parallel matching calculation is performed, and a dynamic weight allocation mechanism is designed at the same time, the normalized platform ID and the cross-platform mapping relationship table are output, and a matching confidence report is generated;
[0041] Initialize the multi-algorithm fusion engine and configure the parallel computing framework;
[0042] Input each feature vector in the multidimensional feature vector set into the engine, extract the corresponding formula in the feature similarity calculation rule library, and perform parallel matching calculation;
[0043] Real-time monitoring of data quality on each platform;
[0044] When the platform data quality is low, the weights are automatically adjusted through the online feedback network, and the final matching score is calculated using the weighted average formula;
[0045] Build a decision-making model based on reinforcement learning and use historical matching data to train the decision-making model based on reinforcement learning;
[0046] When there is a conflict in the matching results, the conflict resolution mechanism is triggered to automatically resolve the conflict;
[0047] For conflicts that cannot be resolved automatically, generate an audit report and record the cause of the conflict;
[0048] For successfully matched stores, generate a normalized store ID for each store and record its original ID on different platforms;
[0049] Based on the normalized store ID and the original ID of each store on different platforms, a cross-platform mapping relationship table is constructed, and a matching confidence report is generated.
[0050] As a preferred solution of the multi-platform store normalization method of the present invention, wherein: storing the normalization result in a distributed database based on blockchain includes the following steps:
[0051] Build a distributed database based on blockchain, configure blockchain network and design smart contracts;
[0052] Each store ID is stored as a main chain record, and the mapping relationship of each platform is stored as a side chain;
[0053] Generate a main chain record for each normalized store ID, recording its unique ID and creation timestamp;
[0054] Generate side chain records for the mapping relationship of each platform, recording the platform ID, matching score and mapping timestamp;
[0055] Each time the data is updated, a new version number is generated and a version change log is recorded.
[0056] As a preferred solution of the multi-platform store normalization method of the present invention, wherein: designing an incremental synchronization mechanism, storing data by geographical area, and dynamically allocating node loads through a consistent hashing algorithm, obtaining an updated normalized platform database includes the following steps:
[0057] Set up an event-driven architecture, use Apache Kafka as the message queue, configure producers and consumers, producers are responsible for publishing platform data update events, and consumers are responsible for processing events and triggering incremental synchronization processes;
[0058] Define event types and configure event processing rules;
[0059] When the corresponding platform data is updated, the lightweight re-matching process is triggered through the event-driven architecture;
[0060] Extract the changed fields and recalculate the feature vector;
[0061] Use a multi-algorithm fusion engine to perform local matching calculations on the changed fields, update the matching results, and store the change records on the chain;
[0062] Divide the blockchain network into multiple shards by geographical region, with each shard responsible for storing store data in a specific region;
[0063] Use the consistent hashing algorithm to dynamically distribute node load and obtain updated normalized data.
[0064] In a second aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the multi-platform store unification method described in the first aspect of the present invention is implemented.
[0065] In a third aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the multi-platform store unification method as described in the first aspect of the present invention.
[0066] The beneficial effects of the present invention are as follows: by adopting a hybrid cleaning strategy based on a rule engine and a deep learning model, intelligent standardized processing of raw data from multi-platform stores is achieved, solving the problems of recognition difficulties, data redundancy and low operational efficiency; a unified API interface adaptation layer combined with a distributed crawler system ensures the real-time and integrity of data capture; regular expressions and a Seq2Seq model based on a Transformer architecture are used to effectively improve the standardization level of address and name fields and enhance data consistency; through multimodal feature extraction technology and a multi-algorithm fusion engine, in-depth mining of standardized data sets is achieved, and an accurate set of feature vectors is generated, significantly improving the accuracy of cross-platform data matching; a distributed database built with the help of blockchain technology not only enhances the security and immutability of data, but also optimizes data storage and management efficiency through an incremental synchronization mechanism. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0068] Figure 1 This is a flow chart of the multi-platform store unification method in Example 1.
[0069] Figure 2 Schematic diagram of multimodal feature extraction in Example 1. DETAILED DESCRIPTION
[0070] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0071] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0072] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0073] Example 1, reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a method for normalizing multi-platform stores, including the following steps:
[0074] S1. Capture the original store data of multiple platforms in real time through the API interface, adopt a hybrid cleaning strategy based on rule engine and deep learning model, perform intelligent standardization processing on the original store data of multiple platforms, and obtain a standardized data set.
[0075] S1.1.Multi-platforms include food delivery platforms, navigation platforms, and social media, etc.
[0076] S1.2. Capture the original store data of multiple platforms in real time through the API interface.
[0077] The specific steps are as follows:
[0078] For different platforms, a unified API interface adaptation layer is set up. For platforms that do not support API, a distributed crawler system based on the Scrapy framework is adopted, and dynamic IP proxy and request frequency control are configured; store names, addresses, contact information, business scope and business hours are captured in real time, and the data source platform, capture timestamp and data version number are recorded, and the original data table is output.
[0079] It should be noted that obtaining raw store data from multiple platforms through a unified API interface adaptation layer and a distributed crawler system improves the flexibility and efficiency of data acquisition and ensures the real-time and integrity of the data; the use of a distributed crawler system enhances the data acquisition capabilities for non-API supported platforms, while improving the stability and concealment of the system through dynamic IP proxy and request frequency control.
[0080] Preferably, using a dynamic IP proxy can effectively avoid being blocked by the target website, while request frequency control can help reduce the pressure on the target server and increase the success rate of data capture.
[0081] S1.3. A hybrid cleaning strategy based on a rule engine and a deep learning model is used to perform intelligent standardization processing on the original store data of the multiple platforms to obtain a standardized data set.
[0082] S1.3.1. Use regular expressions to remove extra spaces and punctuation marks in the original data table (for example, clean "Starbucks (XX Plaza Store)" to "Starbucks XX Plaza Store"), and standardize common abbreviations in the address (for example: replace "XX Rd" with "XX Road") to obtain a preliminary cleaned data table.
[0083] Furthermore, the cleaned data needs to be quality checked to identify and mark abnormal data. Specific inspection rules include: the name field is empty or too short (such as a length less than 2), the address field cannot be parsed to obtain information at the provincial, municipal, or district level, the contact field does not meet format requirements (such as a mobile phone number that is not 11 digits long), and the business hours field format is incorrect (such as containing illegal characters). For abnormal data, mark its status as "pending review" and record the type of abnormality (such as "missing name" or "invalid address").
[0084] S1.3.2. Intelligently convert the multi-language and multi-format data in the data table after preliminary cleaning, extract store-related corpus from the public multi-language parallel corpus, and combine it with the custom industry corpus (such as the Chinese and English store names captured from platforms such as Meituan and Ele.me) to perform sub-word segmentation, remove stop words, and unify the case of the corpus.
[0085] It should be noted that through intelligent conversion of multi-language and multi-format data, the ability to understand multi-language and multi-format data is enhanced, and the consistency and availability of data are improved. Traditional methods often rely on fixed translation rules and are difficult to adapt to the ever-changing language environment.
[0086] S1.3.3. Using the Seq2Seq model based on the Transformer architecture, the original name and address are input into the encoder, and the standardized Chinese name and address are output through the decoder; the standardized address is input into the geocoding service, converted into longitude and latitude coordinates, and the detailed address hierarchy information is parsed; the multi-language and multi-format data that has undergone preliminary cleaning and intelligent conversion, as well as the address geocoding and hierarchical parsing data are integrated to form a standardized data set.
[0087] To further explain, the Seq2Seq model (sequence to sequence model) is a deep learning architecture that is mainly used to handle sequence to sequence tasks. This model usually consists of two main parts: encoder and decoder. Both parts are usually recurrent neural networks (RNN), but other types of networks such as Transformer can also be used.
[0088] The encoder converts the input sequence into a fixed-length vector that contains information about the input sequence. In the original Seq2Seq model, the encoder is an RNN that reads the elements in the input sequence one by one and updates its internal state. Ultimately, the encoder's state is used as a representation of the entire input sequence and passed to the decoder; the decoder receives a fixed-length vector from the encoder and generates an output sequence based on it. Similarly, in the original design, the decoder is also an RNN that starts with the vector received by the encoder and gradually generates each element of the output sequence. When generating each output, the decoder also takes into account all previously generated elements.
[0089] Furthermore, in order to overcome the problem that fixed-length vectors may not fully capture the information of long sequences, the attention mechanism was introduced into the Seq2Seq model. Through the attention mechanism, the decoder can "focus on" different parts of the input sequence when generating each output element, instead of relying solely on a single fixed-length vector, enabling the model to process long sequences more effectively and improve the performance of tasks such as translation.
[0090] Seq2Seq models are widely used in tasks that require conversion of sequence data, such as:
[0091] Machine Translation: Converting sentences from one language into another.
[0092] Speech Recognition: Converting speech signals into text.
[0093] Text summarization: Extract key information from longer documents and generate short summaries.
[0094] Question-answering system: Generate answers based on questions.
[0095] Chatbots: Generate responses based on user input.
[0096] S2. Extract the standardized data set through multimodal feature extraction technology to generate a multidimensional feature vector set.
[0097] S2.1. Use the pre-trained BERT multilingual model to extract semantic features of the name field in the standardized dataset; input the name field into the BERT model to obtain the contextual embedding vector of each character; use the bidirectional LSTM network to further extract the contextual semantic features of the name and generate a name semantic vector.
[0098] Furthermore, the pre-trained BERT multilingual model obtains its initial weights through unsupervised learning on a large amount of text data. The pre-training process of BERT mainly relies on two tasks: masked language model and next sentence prediction, as follows:
[0099] The Masked Language Model (MLM) randomly masks 15% of the words in the input text (i.e., characters or sub-word units in a multilingual environment), and then the model needs to predict these masked words based on the context. This approach allows the model to learn the meaning of a word in its context because the model must infer the masked words based on the surrounding words.
[0100] Next Sentence Prediction (NSP) is a method that not only learns the lexical relationship within a single sentence, but also learns the relationship between sentences. It accepts paired sentences and needs to determine whether the second sentence is the sentence that immediately follows the first sentence in the original text, which helps the model understand longer-range language structures and the relationship between paragraphs.
[0101] Preferably, traditional methods may only rely on simple feature extraction methods such as bag-of-words model or TF-IDF, which cannot effectively capture the semantic relationship between words. The present invention uses the context understanding ability of BERT and the memory function of LSTM to make the model more accurate when processing names with complex grammatical structures or semantics, and can capture the deep semantic information of the text, while the bidirectional LSTM can better understand the order and dependency of the text, thereby enhancing the feature expression ability.
[0102] S2.2. Extract spatial features from the longitude and latitude coordinates of the address field in the standardized data set; use the Geohash algorithm to encode the longitude and latitude coordinates into a string; convert the Geohash string into a binary code to generate a spatial vector.
[0103] To further illustrate, traditional geocoding methods may directly use longitude and latitude values as features, but this is not convenient for comparing the relative position relationship between different locations. The use of the Geohash algorithm can effectively compress geographic location information into a short string and retain a certain degree of geographic proximity, which is convenient for subsequent calculation of similarity.
[0104] S2.3. Extract the unique identifier of the contact information field in the standardized data set, use a regular expression to extract part of the contact information as the unique identifier; perform hash encoding on the unique identifier to generate a contact information vector.
[0105] Specifically, some numbers in the contact information use the first 7 digits as a unique identifier. The first 7 digits of a mobile phone number can determine the location and service provider, but do not contain personal specific information. This ensures a certain degree of uniqueness and reduces the risk of directly exposing complete personal information.
[0106] It should be noted that the use of regular expressions can quickly identify and distinguish different contact methods while ensuring privacy protection, and the selection of appropriate hash functions ensures a low collision rate while maintaining computational efficiency.
[0107] S2.4. Use the TextCNN model based on the attention mechanism to extract keyword features of the business scope in the standardized data set; input the business scope text into the TextCNN model, which includes an embedding layer, a convolution layer, an attention layer, and a fully connected layer; the convolution layer uses multiple convolution kernels of different sizes to extract local features, and the attention layer calculates the importance weight of each word, and the output is the keyword vector of the business scope.
[0108] It is further explained that TextCNN combined with the attention mechanism can not only capture the key information in the text, but also highlight more important words and improve the effectiveness of feature expression; by adjusting the convolution kernel size to adapt to the requirements of different types of text lengths.
[0109] S2.5. Integrate the name semantic vector, space vector, contact information vector and keyword vector into a multi-dimensional feature vector set for each store, and design a feature similarity calculation rule base.
[0110] It should be noted that this comprehensive feature representation method can more comprehensively reflect the characteristics of a store, help improve the accuracy of cross-platform matching, adjust the weight of each feature vector according to the specific application scenario, and optimize the effect of the overall feature representation.
[0111] S2.5.1. The specific steps for designing feature similarity calculation rules are as follows:
[0112] The semantic vectors of the names of the two stores are extracted from the multidimensional feature vector set, and the cosine similarity is used to measure the semantic similarity of the two store names. The expression is:
[0113] ;
[0114] in, Indicates the semantic similarity of the name, and Represent the name semantic vectors of the two stores respectively.
[0115] Set a name similarity threshold. When the cosine similarity is higher than the name similarity threshold, the names are considered to match. When the cosine similarity is lower than the name similarity threshold, the names are considered to not match and are marked as "pending further review."
[0116] Extract the Geohash codes of the two stores and measure the spatial proximity of the two store addresses by calculating the Geohash prefix overlap, which is expressed as:
[0117] ;
[0118] in, Indicates the spatial similarity calculated based on geographic information (such as the Geohash code of the address). Indicates the common prefix length between the Geohash codes of two addresses. and They are the total length of the Geohash codes of the two addresses respectively.
[0119] Set the overlap threshold. When the overlap is higher than the overlap threshold, the addresses are considered to match. When the overlap is lower than the overlap threshold, the two store addresses are considered to be mismatched and can be marked as "pending manual review". Extract the contact information vectors of the two stores. For mobile phone numbers, use the exact match method to match. If the first 7 digits are exactly the same, the similarity is 1.0; otherwise, it is 0.0. For other contact information, use fuzzy hash matching to calculate the Hamming distance between the two vectors. The expression is:
[0120] ;
[0121] in, represents the Hamming distance, and Represent two binary vectors of equal length, Represents the total length of the elements of the binary vector, that is, the number of bits contained in each vector, represents the element index of a binary vector, Represents the first elements, Represents the first elements, Represents a bitwise XOR operation, returning 1 if and only if the corresponding two bits are different, otherwise it returns 0.
[0122] Set a contact information similarity threshold. When the calculated Hamming distance is lower than the contact information similarity threshold, the contact information is considered to match. When the Hamming distance is higher than the set contact information similarity threshold, it is considered to be unmatched, and whether to conduct further manual review is determined based on business needs.
[0123] Extract the keyword vectors of the two stores, and measure the similarity of the business scope of the two stores by calculating the cosine similarity of the keyword vectors. The formula is the same as the expression of semantic similarity.
[0124] A business scope similarity threshold is set. When the cosine similarity of the keyword vector is higher than the business scope similarity threshold, the business scope is determined to be matched. When the similarity is lower than the set business scope similarity threshold, it is considered to be mismatched, and additional data verification is required to confirm whether it is really mismatched. The name similarity, address similarity, contact information similarity and business scope similarity calculation rules are integrated into a similarity calculation rule library.
[0125] To further illustrate, the similarity calculation rule library includes a rule ID, a similarity calculation formula, a similarity threshold, and a rule description (eg, "name similarity calculation rule").
[0126] S3. Input the multi-dimensional feature vector set into the multi-algorithm fusion engine, perform parallel matching calculations, design a dynamic weight allocation mechanism, output the normalized platform ID and cross-platform mapping relationship table, and generate a matching confidence report.
[0127] S3.1. Initialize the multi-algorithm fusion engine and configure the parallel computing framework; input each feature vector in the multi-dimensional feature vector set into the engine, extract the corresponding formula in the feature similarity calculation rule library, and perform parallel matching calculation.
[0128] It should be noted that the multi-algorithm fusion engine can support multiple matching algorithms (such as rule-based matching, machine learning models, etc.), and improve processing speed and cope with large-scale data sets through a parallel computing framework.
[0129] S3.2. Monitor the data quality of each platform in real time; when the platform data quality is low (such as the address missing rate exceeds 30%), automatically adjust the weight through the online feedback network; for example, increase the name weight to 60%, reduce the address weight to 20%, and use the weighted average formula to calculate the final matching score.
[0130] It is further explained that the use of streaming processing frameworks such as Apache Kafka for real-time monitoring and feedback can promptly detect and correct low-quality data, ensuring the accuracy of the subsequent matching process. However, most existing systems lack automated data quality monitoring mechanisms and often require manual intervention and inspection.
[0131] S3.3. Build a decision model based on reinforcement learning, and use historical matching data to train the decision model of reinforcement learning; when there is a conflict in the matching results, trigger the conflict resolution mechanism; submit the conflicting data to the manual review queue, and the reviewer determines whether it is the same store based on the matching score and historical data; for conflicts that cannot be resolved automatically, generate an audit report and record the cause of the conflict; for stores that are successfully matched, generate a normalized store ID for each store, and record its original ID on different platforms; based on the normalized store ID and the original ID of each store on different platforms, build a cross-platform mapping relationship table, and generate a matching confidence report.
[0132] To further explain, the conflict resolution mechanism refers to a series of measures taken by the system to resolve conflicts that are difficult to resolve automatically during the multi-platform store information matching process. The specific steps are as follows:
[0133] Use a trained reinforcement learning model (such as DQN) to predict the best solution based on historical data. Assuming that the conflict is caused by inconsistent addresses, the model will adjust the weights based on the similarity score and recalculate the final matching score. If the reinforcement learning model cannot solve the problem, the conflict resolution mechanism is triggered.
[0134] Conflict resolution mechanism: Determine the specific cause of the conflict and obtain more information from other data sources or historical records to help resolve the conflict; generate multiple possible solutions based on the collected information; submit conflicts that cannot be resolved automatically to the manual review queue, and for each conflict and its solution, record the cause of the conflict and the resolution process in detail.
[0135] S4. Store the normalized results in a distributed database based on blockchain, design an incremental synchronization mechanism, store the data by geographical region, and dynamically distribute the node load through the consistent hashing algorithm to obtain an updated normalized platform database.
[0136] S4.1. Build a distributed database based on blockchain, configure the blockchain network and design smart contracts; store each store ID as a main chain record and each platform mapping relationship as a side chain; generate a main chain record for each normalized store ID, record its unique ID and creation timestamp; generate a side chain record for the mapping relationship of each platform, record the platform ID, matching score and mapping timestamp; each time the data is updated, generate a new version number and record the version change log.
[0137] It should be noted that storing each store ID as a record on the main chain and recording its unique ID and creation timestamp provides an unalterable data storage mechanism, enhancing the integrity and reliability of the data; through the side chain mechanism, the mapping relationship between different platforms can be flexibly managed and expanded while maintaining data consistency.
[0138] S4.2. Set up an event-driven architecture to monitor platform data updates in real time and trigger incremental synchronization; use Apache Kafka is used as a message queue, and producers and consumers are configured. Producers are responsible for publishing platform data update events (such as "store name update"), and consumers are responsible for processing events and triggering incremental synchronization processes; define event types (such as "name update", "address update", and "contact information update"), and configure event processing rules; for example, when a "name update" event is received, the name field is triggered to re-match; when the corresponding platform data is updated, a lightweight re-matching process is triggered through the event-driven architecture; extract the changed fields (such as name, address), and recalculate the feature vector; for example, the name field is re-entered into the BERT model to generate a semantic vector, and the address field is recalculated to calculate the Geohash code; use a multi-algorithm fusion engine to perform local matching calculations on the changed fields instead of full processing, for example, only recalculate the name similarity and address similarity, update the matching results, and store the change records on the chain; divide the blockchain network into multiple shards according to geographical regions (such as North China, East China, and South China), and each shard is responsible for storing store data in a specific area; use a consistent hashing algorithm to dynamically distribute node loads to obtain updated normalized data.
[0139] It should be noted that by defining event types and configuring corresponding processing rules, fine-grained data updates and synchronization are achieved, reducing the need for full processing and improving efficiency; unnecessary calculations are reduced through local updates, improving the system's response speed and resource utilization; dividing the blockchain network into multiple shards according to geographical regions improves the efficiency of data storage and query, while enhancing the scalability of the system.
[0140] This embodiment also provides a computer device, which is applicable to the multi-platform store unification method, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the multi-platform store unification method proposed in the above embodiment.
[0141] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0142] The present embodiment also provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for realizing the unification of multi-platform stores as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, disk or optical disk.
[0143] In summary, the present invention achieves intelligent standardized processing of raw data of multi-platform stores by adopting a hybrid cleaning strategy based on a rule engine and a deep learning model, solving the problems of difficult identification, data redundancy and low operational efficiency; a unified API interface adaptation layer combined with a distributed crawler system ensures the real-time and integrity of data capture; regular expressions and a Seq2Seq model based on a Transformer architecture are used to effectively improve the standardization level of address and name fields and enhance data consistency; multimodal feature extraction technology and a multi-algorithm fusion engine are used to achieve deep mining of standardized data sets, generate accurate feature vector sets, and significantly improve the accuracy of cross-platform data matching; a distributed database built with the help of blockchain technology not only enhances the security and immutability of data, but also optimizes data storage and management efficiency through an incremental synchronization mechanism.
[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for normalizing stores on multiple platforms, characterized by: include, The original store data of multiple platforms is captured in real time through the API interface, and a hybrid cleaning strategy based on the rule engine and the deep learning model is adopted to perform intelligent standardization processing on the original store data of multiple platforms to obtain a standardized data set; the multiple platforms include food delivery platforms, navigation platforms and social media; The data in the standardized data set is extracted by multimodal feature extraction technology to generate a multidimensional feature vector set, including the following steps: The pre-trained BERT multilingual model is used to extract semantic features from the name field in the standardized dataset; Input the name field into the BERT model to obtain the context embedding vector for each character; Use a bidirectional LSTM network to further extract contextual semantic features of the name field and generate a name semantic vector; Extract spatial features from the latitude and longitude coordinates of the address field in the standardized dataset; Use the Geohash algorithm to encode the longitude and latitude coordinates into a string; Convert the Geohash string into binary code to generate a space vector; Extract unique identifiers from the contact information field in the standardized data set, and use regular expressions to extract partial numbers of the contact information as unique identifiers; Hash the unique identifier to generate a contact vector; The TextCNN model based on the attention mechanism is used to extract keyword features from the business scope in the standardized data set; The business scope text is input into the TextCNN model. The convolution layer uses multiple convolution kernels of different sizes to extract local features. The attention layer calculates the importance weight of each word and outputs the keyword vector of the business scope. Integrate name semantic vector, space vector, contact information vector and keyword vector into a multi-dimensional feature vector set for each store, and design a feature similarity calculation rule base; Input the multi-dimensional feature vector set into the multi-algorithm fusion engine, perform parallel matching calculations, design a dynamic weight allocation mechanism, output the normalized platform ID and cross-platform mapping relationship table, and generate a matching confidence report; The normalized results are stored in a distributed database based on blockchain, an incremental synchronization mechanism is set up, the data is stored by geographical region, and the node load is dynamically distributed through the consistent hashing algorithm to obtain the updated normalized platform data.
2. The method for normalizing multi-platform stores as claimed in claim 1, characterized in that: The real-time acquisition of original store data from multiple platforms through the API interface includes the following steps: Design a unified API interface adaptation layer for different platforms. For platforms that do not support API, use a distributed crawler system based on the Scrapy framework, configure dynamic IP proxy and request frequency control; Capture store name, address, contact information, business scope and business hours in real time, record the data source platform, capture timestamp and data version number, and output the original data table.
3. The method for normalizing multi-platform stores as claimed in claim 2, characterized in that: A hybrid cleaning strategy based on rule engine and deep learning model is used to perform intelligent standardization on the original store data of multiple platforms to obtain a standardized data set, including the following steps: Use regular expressions to remove extra spaces and punctuation marks in the original data table and standardize common abbreviations in the address to obtain a preliminary cleaned data table; Intelligently convert the multi-language and multi-format data in the data table after preliminary cleaning, extract store-related corpus from the public multi-language parallel corpus, and combine it with the custom industry corpus to perform sub-word segmentation, remove stop words, and unify the capitalization of the corpus; A Seq2Seq model based on the Transformer architecture is used to input the original name and address into the encoder and output the standardized Chinese name and address through the decoder; Input the standardized address into the geocoding service, convert it into longitude and latitude coordinates, and parse out the detailed address level information; Integrate multi-language and multi-format data that have undergone preliminary cleaning and intelligent conversion, as well as data on address geocoding and hierarchical parsing, to form a standardized data set.
4. The method for normalizing multi-platform stores as claimed in claim 3, characterized in that: The design feature similarity calculation rule includes the following steps: Extract the semantic vectors of the names of the two stores from the multidimensional feature vector set, and use cosine similarity to measure the semantic similarity of the two store names; Extract the Geohash codes of the two stores and measure the spatial proximity of the two store addresses by calculating the Geohash prefix overlap; Extract the contact information vectors of the two stores and use the exact matching method to match the mobile phone numbers; For other contact information, the Hamming distance between two contact information vectors is calculated using fuzzy hash matching; Extract the keyword vectors of the two stores and measure the similarity of the business scope of the two stores by calculating the cosine similarity of the keyword vectors; The name similarity, address similarity, contact information similarity and business scope similarity calculation rules are integrated into a similarity calculation rule library.
5. The method for normalizing multi-platform stores as claimed in claim 4, characterized in that: Input the multi-dimensional feature vector set into the multi-algorithm fusion engine, perform parallel matching calculations, design a dynamic weight allocation mechanism, output the normalized platform ID and cross-platform mapping relationship table, and generate a matching confidence report; Initialize the multi-algorithm fusion engine and configure the parallel computing framework; Input each feature vector in the multidimensional feature vector set into the engine, extract the corresponding formula in the feature similarity calculation rule library, and perform parallel matching calculation; Real-time monitoring of data quality on each platform; When the platform data quality is low, the weights are automatically adjusted through the online feedback network, and the final matching score is calculated using the weighted average formula; Build a decision-making model based on reinforcement learning and use historical matching data to train the decision-making model based on reinforcement learning; When there is a conflict in the matching results, the conflict resolution mechanism is triggered to automatically resolve the conflict; For conflicts that cannot be resolved automatically, generate an audit report and record the cause of the conflict; For successfully matched stores, generate a normalized store ID for each store and record its original ID on different platforms; Based on the normalized store ID and the original ID of each store on different platforms, a cross-platform mapping relationship table is constructed, and a matching confidence report is generated.
6. The method for normalizing multi-platform stores as claimed in claim 5, characterized in that: Storing the normalized results in a distributed database based on blockchain includes the following steps: Build a distributed database based on blockchain, configure blockchain network and design smart contracts; Each store ID is stored as a main chain record, and the mapping relationship of each platform is stored as a side chain; Generate a main chain record for each normalized store ID, recording its unique ID and creation timestamp; Generate side chain records for the mapping relationship of each platform, recording the platform ID, matching score and mapping timestamp; Each time the data is updated, a new version number is generated and a version change log is recorded.
7. The method for normalizing multiple platforms and stores as claimed in claim 6, characterized in that: Design an incremental synchronization mechanism to store data by geographical region, and dynamically distribute node loads through a consistent hashing algorithm to obtain an updated normalized platform database. The following steps are included: Set up an event-driven architecture, use Apache Kafka as the message queue, configure producers and consumers, producers are responsible for publishing platform data update events, and consumers are responsible for processing events and triggering incremental synchronization processes; Define event types and configure event processing rules; When the corresponding platform data is updated, the lightweight re-matching process is triggered through the event-driven architecture; Extract the changed fields and recalculate the feature vector; Use a multi-algorithm fusion engine to perform local matching calculations on the changed fields, update the matching results, and store the change records on the chain; Divide the blockchain network into multiple shards by geographical region, with each shard responsible for storing store data in a specific region; Use the consistent hashing algorithm to dynamically distribute node load and obtain updated normalized data.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the multi-platform store unification method described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multi-platform store unification method described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Offline new retail store passenger flow multi-attribute single-model identification method
CN110991528A
Text extraction method and device for multi-modal data, refrigeration equipment and medium
CN116956209A