Special work information automatic screening system for shopping center management personnel

Through the automatic screening system for special information for shopping center managers, the problems of inconsistent data quality and untargeted recommendations in traditional systems are solved, and high-quality data screening and personalized recommendations are realized to meet the needs of managers.

CN120336624APending Publication Date: 2025-07-18BEIJING INNOVATION CHINA BUSINESS UNITED BUSINESS MANAGEMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510383545.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional information screening systems are difficult to effectively filter and recommend content that meets user interests and needs in massive data, and the data quality is uneven, making it difficult for managers to obtain the most relevant and useful information.

Method used

A special information automatic screening system for shopping center managers is designed, including data collection, preprocessing, fusion, filtering and personalized recommendation modules. Through data cleaning, noise removal, unified format, quality evaluation and multi-level screening, we ensure high quality and personalized recommendation of data.

Benefits of technology

It improves the accuracy and pertinence of information screening, ensures that managers obtain the most relevant and useful information, and the system has good maintainability and scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336624A_ABST
    Figure CN120336624A_ABST
Patent Text Reader

Abstract

The invention, which belongs to the technical field of information management, discloses an automatic screening system for special work information of shopping center managers, comprising a data acquisition module, a data preprocessing module, a data fusion module, an information filtering module and a personalized recommendation module. The data acquisition module is used for continuously acquiring sales data, inventory data and customer information from an internal management system, a monitoring system and an external public data source of a shopping center, and outputting the acquired data according to a preset format; the data preprocessing module is used for performing cleaning, noise removal, data format unification and standardization processing on the original data output by the data acquisition module to obtain preprocessed data; according to the method, integrity, accuracy and consistency evaluation is carried out on the data of the data sources, the data fusion rule is automatically adjusted according to the evaluation result, the high quality of the fused data set is ensured, and the measures jointly improve the reliability of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information management, and more specifically, to an automatic information screening system dedicated to the work of shopping mall managers. Background Art

[0002] In the digital age of information overload, users are faced with a large number of information choices, and it has become difficult for users to screen out valuable content from the vast amount of information. How to effectively filter and recommend content that meets the interests and needs of users has become a key issue. Recommendation systems have emerged, aiming to provide personalized information recommendations by analyzing user behavior and preferences. Each user has different interests and needs, and the traditional "one-size-fits-all" information push method can no longer meet the personalized needs. Personalized recommendation systems provide customized content by analyzing user historical behavior and preferences, improving user satisfaction and stickiness. With the progress of machine learning, data mining, and artificial intelligence technologies, it has become possible to build efficient and accurate recommendation systems. These technologies can process massive amounts of data, discover users' potential interests, and improve the quality of recommendations.

[0003] The keyword matching of traditional screening systems is too simple, easily screening out irrelevant information or missing important information, and the data used for statistics is uneven. Many useless data will affect the recommendation results during the recommendation process, lacking pertinence, and it is difficult for managers to obtain the most relevant and useful information. Summary of the Invention

[0004] The present invention proposes an automatic information screening system dedicated to the work of shopping mall managers to solve the above problems.

[0005] Technical Solution: An automatic information screening system dedicated to the work of shopping mall managers includes a data collection module, a data preprocessing module, a data fusion module, an information filtering module, and a personalized recommendation module;

[0006] Data collection module: Continuously collect sales data, inventory data, and customer information from the internal management system, monitoring system, and external public data sources of the shopping mall, and output the collected data in a predetermined format;

[0007] Data preprocessing module: Clean, remove noise, unify the data format, and standardize the original data output by the data collection module to obtain preprocessed data. Then, divide the preprocessed data into different preprocessed data sets according to different brand categories, and then transfer it to the data fusion module;

[0008] Data fusion module: It includes a data quality assessment unit and a fusion unit. The data quality assessment unit calculates the integrity score, accuracy score, and consistency score for different preprocessed data sets. The fusion unit merges the preprocessed data from each data source into a unified fused data set according to the fusion rules;

[0009] Information filtering module: It uses a preliminary screening rule to screen the fused data set output by the data fusion module to obtain the preliminarily screened data, and uses a pre-determined dynamic threshold to perform secondary filtering on the preliminarily screened data, so as to obtain the information set after secondary filtering;

[0010] Personalized recommendation module: Using the information set obtained by the information filtering module, it sorts and outputs the brands corresponding to the brand attributes that meet the user's needs according to the user's input requirements.

[0011] Preferably, the data quality assessment unit in the data fusion module also evaluates the integrity, accuracy, and consistency of the preprocessed data corresponding to the data collected from each data source.

[0012] Preferably, the data collection module, data preprocessing module, data fusion module, information filtering module, and personalized recommendation module adopt a modular design, and data transfer is realized between modules through a standardized interface.

[0013] Preferably, the data quality assessment unit calculates the integrity score, accuracy score, and consistency score as follows:

[0014]

[0015]

[0016] Among them, C is the integrity score, A is the accuracy score, and I is the consistency score; N total is the total number of data items in the preprocessed data set, N miss is the number of missing data items in the preprocessed data set, N error is the number of data items that do not meet the expectations or are incorrect in the preprocessed data set, N compare is the total number of pairs that need to be compared across data sources in the preprocessed data set, N incons is the number of inconsistent data pairs found during the comparison of the preprocessed data set.

[0017] The overall data quality score Q of the preprocessed data is calculated using the weighted average method:

[0018] Q = w1×C + w2×A + w3×I

[0019] Among them, w1, w2, and w3 are the weights of integrity, accuracy, and consistency respectively, designed for users, and satisfy:

[0020] w1 + w2 + w3 = 1.

[0021] Preferably, the working steps of the fusion unit are as follows:

[0022] Step 1: Identify all records of the same brand according to the brand matching rule;

[0023] Step 2: For each attribute of the same brand, merge the pre-processed data respectively according to the method in the attribute merging rule;

[0024] Step 3: For the conflicts found in the case of multiple different values in the same numerical attribute during the merging process, perform conflict handling according to the conflict handling rule, and generate records to be reviewed;

[0025] Step 4: Integrate the data without anomalies and the data after review into a fusion data set.

[0026] Preferably, the brand matching rule is: for the sales data, inventory data, and customer information of this brand, if it contains a unique identifier, then when the keywords recorded in different data sources are exactly the same, it is determined that the record belongs to the same brand; for records lacking a unique identifier, calculate the string similarity of the key text fields in the record, set a threshold of 0.8, and when the similarity is higher than this threshold, it is determined that the two records describe the same brand.

[0027] Preferably, for each attribute of the same brand, the attribute merging rule adopts the following merging method:

[0028] For the numerical attributes of the brand's sales data, inventory data, and customer information, the quality score Q i , for a certain numerical attribute in the sales data, inventory data, and customer information, calculate the weighted average value, and the merged numerical attribute V 合并 The formula is:

[0029]

[0030] where m is the number of data sources providing the attribute data, and V i is the value of this attribute in the i-th data source.

[0031] Preferably, for the case of multiple different values in the same numerical attribute, the conflict handling rule compares the quality scores of each data source, and selects the data value with the highest quality score as the final value; when the quality scores of multiple data sources are the same, mark this attribute as the conflict state and generate a conflict record for subsequent manual verification.

[0032] Preferably, the initial screening rules are as follows: Check the integrity of the data in the fusion dataset. If there are missing key fields, the data is regarded as incomplete and excluded. Then, perform data accuracy verification. For numerical fields, check whether they are within a reasonable range, and records outside the reasonable range will be excluded. Finally, perform data duplication detection to detect and remove duplicate records.

[0033] Preferably, the secondary filtering content is as follows: Select the lower quartile of the quality score as the threshold, filter the records with a comprehensive quality score lower than the dynamic threshold, and retain high-quality data.

[0034] Compared with the prior art, the advantages of the present invention are as follows:

[0035] (1) Through the data preprocessing module, the present invention cleans, denoises, unifies the format, and standardizes the units of the original data, ensuring the integrity and consistency of the data. The data quality assessment unit further assesses the integrity, accuracy, and consistency of the data from each data source, and automatically adjusts the data fusion rules according to the assessment results to ensure the high quality of the fusion dataset. These measures together improve the reliability of the data.

[0036] (2) Through the information filtering module, the present invention adopts an initial screening rule and a secondary filtering mechanism with a dynamic threshold to precisely screen the fusion dataset, excluding incomplete, inaccurate, inconsistent, and duplicate data, ensuring that the retained data has high quality and high relevance. The personalized recommendation module sorts and outputs the screened information set according to the specific needs of the management personnel, ensuring the pertinence and practicality of the recommendation results. This multi-level screening and recommendation mechanism improves the accuracy and pertinence of information screening, ensuring that the management personnel obtain the most relevant and useful information.

[0037] (3) The present invention adopts a modular design, and data transfer and function complementarity are realized between modules through standardized interfaces. This design method makes the system have good maintainability and scalability, facilitating the subsequent upgrade and expansion of functions to meet the changing needs of shopping mall management personnel. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is the overall system module diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] Example: Refer to Figure 1 , an automatic information screening system for shopping mall management personnel, including a data acquisition module, a data preprocessing module, a data fusion module, an information filtering module, and a personalized recommendation module;

[0040] Data collection module: Continuously collect sales data, inventory data, and customer information from the shopping mall's internal management system, monitoring system, and external public data sources, and output the collected data in a predetermined format;

[0041] Data preprocessing module: Clean, remove noise, unify the data format, and standardize the original data output by the data collection module to obtain preprocessed data. Then, divide the preprocessed data into different preprocessed data sets according to different brand categories and transfer them to the data fusion module;

[0042] It includes the following steps:

[0043] Data cleaning: Clean the original data, delete or correct missing values, duplicate values, and outliers. For missing data, choose to delete the records containing missing values according to business requirements, or use methods such as mean, median, and mode to fill them; Detect and delete duplicate records to ensure data uniqueness; Identify outliers based on business rules and choose to correct or delete them according to specific situations.

[0044] Noise removal: Use techniques such as filtering and smoothing to remove random noise in the data and enhance data reliability. For time series data, use methods such as moving average or Kalman filtering to smooth the data; For spatial data, apply methods such as Gaussian filtering or median filtering to remove noise.

[0045] Data format unification: Convert the original data from different sources and different formats into a unified format to ensure data consistency; Unify various date and time representation methods into the ISO 8601 standard format; Convert strings in different encoding formats into UTF-8 encoding.

[0046] Data unit standardization: Convert data with the same meaning but different units into a unified unit to ensure data comparability; Convert units such as inches and feet into meters; Convert units such as pounds and ounces into kilograms; Convert different currencies into a unified currency according to the exchange rate.

[0047] Data verification: Verify the preprocessed data to ensure data integrity, accuracy, and consistency. Check whether there are missing or incomplete situations in the data, and check whether the data conforms to business rules or expected ranges to ensure that the data is consistent among different sources or different records.

[0048] Data transfer: Transfer the data preprocessed as above to the data fusion module for subsequent data fusion and analysis.

[0049] Data fusion module: It includes a data quality assessment unit and a fusion unit. The data quality assessment unit calculates the integrity score, accuracy score, and consistency score for different preprocessed data sets. The fusion unit merges the preprocessed data from each data source into a unified fused data set according to the fusion rules;

[0050] Information filtering module: It uses a preliminary screening rule to screen the fused data set output by the data fusion module to obtain the preliminarily screened data, and uses a pre-determined dynamic threshold to perform secondary filtering on the preliminarily screened data, so as to obtain the information set after secondary filtering;

[0051] Personalized recommendation module: Using the information set obtained by the information filtering module, sort and output the brands corresponding to the brand attributes that meet the user's needs according to the user's input requirements

[0052] The data quality assessment unit in the data fusion module evaluates the integrity, accuracy, and consistency of the data collected from each data source.

[0053] Each module adopts a modular design, and data transfer and functional complementarity are realized between modules through standardized interfaces.

[0054] The data quality assessment unit calculates the scores for each dimension as follows:

[0055]

[0056] Among them, C is the integrity score, A is the accuracy score, I is the consistency score; N tital is the total number of data items in the preprocessed data set, N miss is the number of missing data items in the preprocessed data set, N error is the number of data items that do not meet the expectations or are incorrect in the preprocessed data set, N compare is the total number of pairs that need to be compared across data sources in the preprocessed data set, N incons is the number of inconsistent data pairs found during the comparison of the preprocessed data set.

[0057] The overall data quality score Q of the preprocessed data is calculated using the weighted average method:

[0058] Q = w1×C + w2×A + w3×I

[0059] Among them, w1, w2, and w3 are the weights of integrity, accuracy, and consistency respectively, designed for the users, and satisfy:

[0060] w1 + w2 + w3 = 1.

[0061] Preferably, the working steps of the fusion unit are as follows:

[0062] Step 1: Identify all records of the same brand according to the brand matching rules;

[0063] Step 2: For each attribute of the same brand, merge the preprocessed data respectively according to the methods in the attribute merging rules;

[0064] Step 3: For the conflicts found in the case of multiple different values in the same numerical attribute during the merging process, perform conflict handling according to the conflict handling rules and generate records to be audited;

[0065] Step 4: Integrate the data without anomalies and the data after auditing into a fusion dataset.

[0066] Preferably, the brand matching rules are as follows: For the sales data, inventory data and customer information of this brand, if it contains a unique identifier, then when the keywords recorded in different data sources are exactly the same, it is determined that this record belongs to the same brand; for records lacking a unique identifier, calculate the string similarity of the key text fields in the record, set a threshold of 0.8, and when the similarity is higher than this threshold, it is determined that the two records describe the same brand.

[0067] Preferably, for each attribute of the same brand, the following merging method is adopted for the attribute merging rules:

[0068] For the numerical attributes of the brand's sales data, inventory data and customer information, the quality score Q of each data source i , for a certain numerical attribute in the sales data, inventory data and customer information, calculate the weighted average value, and the merged numerical attribute V 合并 The formula is:

[0069]

[0070] Among them, m is the number of data sources providing data for this attribute, and V i is the value of this attribute in the i-th data source.

[0071] Preferably, for the case of multiple different values in the same numerical attribute, the conflict handling rules compare the quality scores of each data source and select the data value with the highest quality score as the final value; when the quality scores of multiple data sources are the same, mark this attribute as in a conflict state and generate a conflict record for subsequent manual verification.

[0072] Preferably, the initial screening rules are as follows: Check the integrity of the data in the fusion dataset. If there is a missing key field, the data is regarded as incomplete and excluded; then perform data accuracy verification. For numerical fields, check whether they are within a reasonable range, and records outside the reasonable range will be excluded; finally, perform data duplication detection and detect and remove duplicate records.

[0073] Preferably, the secondary filtering content is as follows: Select the lower quartile of the quality score as the threshold, filter the records with a comprehensive quality score lower than the dynamic threshold, and retain high-quality data.

[0074] The lower quartile (the first quartile, Q1) is an indicator in statistics used to describe the data distribution. It represents the value at the 25% position after arranging all the data in ascending order, that is, 25% of the data is less than or equal to this value. The steps to calculate the lower quartile are as follows:

[0075] 1. Data sorting: Arrange all the data in ascending order.

[0076] 2. Determine the position: Calculate the position of the lower quartile in the sorted data. The formula is:

[0077] L = (n + 1) × 0.25

[0078] where n is the total number of data, and L is the position of the lower quartile.

[0079] 3. Determine the lower quartile:

[0080] If L is an integer, then the L-th data is the lower quartile. If L is not an integer, then the lower quartile is the linear interpolation of the L-th data and the (L + 1)-th data.

[0081] Example:

[0082] Suppose there is the following data set:

[0083] 6, 7, 15, 36, 39, 40, 41, 42, 43, 47, 496, 7, 15, 36, 39, 40, 41, 42, 43, 47, 496, 7, 15, 36, 39, 40, 41, 42, 43, 47, 49

[0084] 1. Data sorting: The data has been arranged in ascending order.

[0085] 2. Determine the position: L = (11 + 1) × 0.25 = 3; Therefore, the lower quartile is located at the 3rd data.

[0086] Determine the lower quartile: The 3rd data is 15, so the lower quartile is 15.

[0087] The foregoing has shown and described the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments, and the above embodiments and the descriptions in the specification are only preferred examples of the present invention, which are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of the present invention claimed is defined by the appended claims and their equivalents.

Claims

1. An automatic information screening system dedicated to the work of shopping center management personnel, characterized in that, It includes a data collection module, a data preprocessing module, a data fusion module, an information filtering module, and a personalized recommendation module; Data collection module: Continuously collect sales data, inventory data, and customer information from the shopping mall's internal management system, monitoring system, and external public data sources, and output the collected data in a predetermined format; Data preprocessing module: Clean, remove noise, unify the data format, and standardize the raw data output by the data collection module to obtain preprocessed data. Then, divide the preprocessed data into different preprocessed data sets according to different brand categories, and transfer them to the data fusion module; Data fusion module: It includes a data quality assessment unit and a fusion unit. The data quality assessment unit calculates the integrity score, accuracy score, and consistency score for different preprocessed data sets. The fusion unit merges the preprocessed data from each data source into a unified fusion data set according to the fusion rules; Information filtering module: Use the initial screening rules to screen the fusion data set output by the data fusion module to obtain the initially screened data, and use a pre-determined dynamic threshold to perform secondary filtering on the initially screened data to obtain the information set after secondary filtering; Personalized recommendation module: Use the information set obtained by the information filtering module to sort and output the brands corresponding to the brand attributes that meet the user's needs according to the user's input requirements.

2. The automatic information screening system for the work of shopping mall management personnel according to claim 1, characterized in that The data quality assessment unit in the data fusion module also evaluates the integrity, accuracy, and consistency of the preprocessed data corresponding to the data collected by each data source.

3. The automatic information screening system dedicated to the work of shopping mall managers according to claim 1, characterized in that, The data collection module, data preprocessing module, data fusion module, information filtering module, and personalized recommendation module adopt a modular design, and data transfer is realized between modules through standardized interfaces.

4. The automatic information screening system for the work of shopping mall management personnel according to claim 1, characterized in that The data quality assessment unit calculates the integrity score, accuracy score, and consistency score as follows: Among them, C is the integrity score, A is the accuracy score, and I is the consistency score; N total is the total number of data items in the preprocessed dataset, N miss is the number of missing data items in the preprocessed dataset, N error is the number of data items that do not meet the expectations or are incorrect in the preprocessed dataset, N compare is the total number of pairs that need to be compared across data sources in the preprocessed dataset, N incons is the number of inconsistent data pairs found during the comparison of the preprocessed dataset; The overall data quality score Q of the preprocessed data is calculated using the weighted average method: Q = w1×C + w2×A + w3×I Among them, w1, w2, and w are the weights of integrity, accuracy, and consistency respectively, designed for users, and satisfy: w1 + w2 + w3 = 1.

5. The automatic information screening system for shopping mall management staff work according to claim 4, characterized in that, The working steps of the fusion unit are as follows: Step 1: Identify all records of the same brand according to the brand matching rules; Step 2: For each attribute of the same brand, merge the preprocessed data respectively according to the method in the attribute merging rules; Step 3: For the conflicts found in the case of multiple different values in the same numerical attribute during the merging process, perform conflict handling according to the conflict handling rules and generate records to be reviewed; Step 4: Integrate the data without abnormalities and the data after review into a fusion data set.

6. The automatic information screening system for the work of shopping mall managers according to claim 4, characterized in that, The brand matching rules are as follows: For the sales data, inventory data, and customer information of this brand, if it contains a unique identifier, when the keywords recorded in different data sources are exactly the same, it is determined that this record belongs to the same brand; for records lacking a unique identifier, calculate the string similarity of the key text fields in the record, set a threshold of 0.8, and when the similarity is higher than this threshold, it is determined that the two records describe the same brand.

7. An automatic information screening system for shopping mall management staff work as claimed in claim 4, characterized in that For each attribute of the same brand, the following merging method is adopted for the attribute merging rule: For the numerical attributes of brand sales data, inventory data, and customer information, the quality score Q of each data source i , for a certain numerical attribute in sales data, inventory data, and customer information, calculate the weighted average, and the combined numerical attribute V 合并 The formula is: where m is the number of data sources providing the attribute data, and V i is the value of the attribute in the i-th data source.

8. The automatic information screening system for shopping mall management staff work according to claim 4, characterized in that, For the case where there are multiple different values in the same numerical attribute, the conflict handling rule compares the quality scores of each data source and selects the data value with the highest quality score as the final value; when the quality scores of multiple data sources are the same, mark this attribute as in a conflict state and generate a conflict record for subsequent manual verification.

9. The automatic information screening system dedicated to the work of shopping mall managers according to claim 1, characterized in that The preliminary screening rule is as follows: check the integrity of the data in the fusion dataset. If there is a missing key field, the data is regarded as incomplete and excluded; then perform data accuracy verification. For numerical fields, check whether they are within a reasonable range, and records outside the reasonable range will be excluded; finally, perform data duplication detection and detect and remove duplicate records.

10. An automatic information screening system dedicated to the work of shopping center managers according to claim 4, characterized in that, The content of the secondary filtering is as follows: select the lower quartile of the quality score as the threshold, filter the records with a comprehensive quality score lower than the dynamic threshold, and retain high-quality data.