Real Estate Data Processing and Warning Method and System Based on Big Data
By extracting, managing and streamlining real estate data and generating risk warnings in combination with real estate attributes, the problem of incomplete analysis in the existing technology is solved, and multi-dimensional real estate information of enterprises is realized, and data quality and database efficiency are improved.
Patent Information
- Application Number
- CN202411541886.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-10-31
AI Technical Summary
In the prior art, corporate risk analysis methods fail to effectively combine real estate data, resulting in incomplete analysis.
By determining the first and second data sources, extracting existing data, performing data governance and streamlining, extracting real estate-related attributes, setting judgment conditions to trigger risk warnings, and building a topology chart to streamline data and fill in missing values.
It realizes multi-dimensional real-time tracking and early warning of enterprise real estate information, ensures data quality, reduces hardware costs, and improves database operation speed.
Smart Images

Figure CN119336818B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a method and system for real estate data processing and early warning based on big data. Background Art
[0002] Big data technology can acquire, store, analyze, and output a large amount of data information. With the help of big data technology, the business data of enterprises can be uniformly sorted out and allocated, enterprise files can be improved, risk identification of enterprises can be carried out, the financial status, credit, and business stability of enterprises can be obtained, and thus help can be provided for the work of relevant personnel.
[0003] In the prior art, a variety of methods have been provided to analyze the situation of enterprises in combination with big data. For example, the Chinese patent document with the publication number CN116468271A discloses a method, system, and medium for enterprise risk analysis based on big data. This method constructs an enterprise risk analysis model through big data, acquires multi-source business data of enterprises, then extracts the characteristic values of the multi-source business data of enterprises to generate multi-source heterogeneous parameters for enterprise operation, and finally inputs the multi-source heterogeneous parameters for enterprise operation into the enterprise risk analysis model to generate risk assessment information. By comparing the risk assessment information with the preset risk standard information, the enterprise risk assessment level is generated. Another example is the Chinese patent document with the publication number CN118153964A, which discloses a method and system for supplier enterprise risk assessment based on big data technology. This method first reads the enterprise data of the supplier enterprise, then selects the corresponding summary model from the pre-trained summary model library according to the enterprise information data, obtains the enterprise operation summary based on the analysis of the enterprise information data and enterprise operation data by the summary model, and then, in response to the enterprise operation data by the risk sniffing model, selects the corresponding feature model from the pre-trained feature model library, obtains the data features of the enterprise operation data based on the feature model, and inputs the data features into the risk assessment model to obtain the risk assessment result of the enterprise.
[0004] However, the above methods only acquire and analyze the business data of enterprises and do not combine real estate data, so there is a problem of incomplete analysis. Summary of the Invention
[0005] To solve the above problems, the present invention provides a method and system for real estate data processing and early warning based on big data to solve the problems in the prior art.
[0006] To achieve the above invention purpose, the present invention proposes a method for real estate data processing and early warning based on big data, including:
[0007] Determine the first data source and the second data source, and perform an overall extraction of the stock data of the first data source and the second data source to obtain the first main data and the first additional data;
[0008] Write the first main data and the first additional data into the database, perform data governance on the first main data in combination with the first additional data to obtain the second main data, and perform data reduction on the first additional data after the governance is completed to obtain the second additional data;
[0009] Set the first attribute, extract the first attribute from the second main data to generate the first form for each enterprise, and extract the second attribute from the first form to generate the second form, where the second attribute includes the real estate mortgage status, the real estate seizure status, the real estate valuation value, and the number of mortgages;
[0010] Extract the first main data and the first additional data again at preset intervals to update the second form in the database;
[0011] Set a judgment condition, and trigger a risk warning when the second attribute in the second form meets the judgment condition.
[0012] Further, the data reduction of the first additional data includes the following steps:
[0013] The second data source includes multiple sub-data sources. Locate the target data of numerical type in the sub-data sources, organize the target data belonging to the same data name in the sub-data sources into a numerical sequence, calculate the first similarity between different data names in different sub-data sources, and divide the data names into multiple first data groups, where the first similarity between any two data names in the first data group is greater than the first threshold;
[0014] Calculate the second similarity between different numerical sequences in the first data group, and construct a topology graph based on the second similarity. The topology graph includes multiple nodes, where the nodes represent the numerical sequences in the first data group, and the optimal similarity is marked between the nodes;
[0015] Select two adjacent nodes in the topology graph, determine the first sequence and the second sequence in the numerical sequences corresponding to the two nodes, and generate a mapping function based on the first sequence and the second sequence. The mapping function is used to transform the first sequence into the second sequence;
[0016] The database deletes the second sequence in the sub-data source, retains the first sequence corresponding to the deleted second sequence and the mapping function to complete the reduction of the first additional data.
[0017] Further, constructing the topology graph includes the following steps:
[0018] Extract two numerical sequences corresponding to the maximum second similarity from the first data group, which are defined as the first initial sequence and the second initial sequence, and the corresponding nodes are respectively defined as the first initial point and the second initial point. Screen the numerical sequences with the maximum second similarity to the first initial sequence or the second initial sequence from the first data group as the third initial sequence, and connect it as the third initial point to the corresponding first initial point or second initial point. Repeat this step until all the numerical sequences in the first data group are extracted, and obtain the topology graph corresponding to the first data group.
[0019] Further, generating the mapping function includes the following steps:
[0020] The numerical sequences corresponding to two adjacent nodes in the topology graph are defined as the first alternative sequence and the second alternative sequence. Generate a first function and a second function based on the first alternative sequence and the second alternative sequence. Use the first function to transform the first alternative sequence into a third alternative sequence, and use the second function to transform the second alternative sequence into a fourth alternative sequence. Calculate the first difference degree between the first alternative sequence and the fourth alternative sequence, and the second difference degree between the second alternative sequence and the third alternative sequence. If the first difference degree is less than the second difference degree, define the first alternative sequence as the first sequence and the first function as the mapping function; otherwise, define the second alternative sequence as the first sequence and the second function as the mapping function.
[0021] Further, after refining the topology graph, if there are a first reserved point and a second reserved point, and all the nodes between the first reserved point and the second reserved point are deleted, define the deleted nodes as target points. Starting from the first reserved point and the second reserved point respectively, calculate the cumulative product value of the reserved points in combination with the optimal similarity. If the cumulative product values of the existing reserved points are all less than the third threshold, restore the reserved point.
[0022] Further, filling in the missing values of the first main data includes the following steps:
[0023] Obtain the first main data with missing values, obtain the key information of the first main data, and search in the first additional data based on the key information to fill in the missing values.
[0024] Further, the data governance includes duplicate data deletion, missing data filling, data transcoding, and data desensitization.
[0025] Further, when the real estate mortgage status changes, the real estate seizure status changes, the change value of the real estate valuation value exceeds the fourth threshold, or the number of mortgages increases, it is determined that the second attribute meets the judgment condition.
[0026] Further, the first additional data extracted again is defined as new data. Whenever the number of the new data reaches the fifth threshold, data reduction is performed on the currently existing new data.
[0027] The present invention also provides a big data-based real estate data processing and early warning system, which is used to implement the above-mentioned big data-based real estate data processing and early warning method. The system includes:
[0028] A collection module that determines a first data source and a second data source. The collection module is used to extract the stock data of the first data source and the second data source as a whole to obtain first main data and first additional data, and extract the first main data and the first additional data again at preset time intervals to update a second form in the database;
[0029] A storage module that writes the first main data and the first additional data into the database, performs data governance on the first main data in combination with the first additional data to obtain second main data, and performs data reduction on the first additional data after the governance is completed to obtain second additional data;
[0030] An extraction module that sets a first attribute, extracts the first attribute from the second main data to generate a first form for each enterprise, and extracts a second attribute from the first form to generate the second form. The second attribute includes the real estate mortgage status, the real estate seizure status, the real estate valuation value, and the number of mortgages;
[0031] An early warning module that sets a judgment condition, and triggers a risk early warning when there is a second attribute in the second form that meets the judgment condition.
[0032] Compared with the prior art, the beneficial effects of the present invention are at least as follows:
[0033] The present invention obtains the real estate data of enterprises through various methods, and selects the data downloaded from the official website as the main data, that is, the first main data. Before analysis, data governance is carried out on the first main data to ensure its data quality. During the governance process, if there are missing data in the main data, the first additional data other than the first main data is used to supplement the missing values, thus ensuring the integrity of the first main data. The present invention enables relevant personnel to quickly obtain the real estate information of enterprises in multiple dimensions, and realizes real-time tracking and early warning of abnormal real estate information of enterprises by setting judgment conditions.
[0034] After the present invention supplements the missing values of the first main data with the first additional data, it also streamlines the first additional data, thereby reducing the occupied space of the first additional data, accelerating the operation speed of the database, and reducing the expenditure cost of hardware. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 is the step flow chart of the real estate data processing and early warning method based on big data of the present invention;
[0036] Figure 2 is the interface diagram of successful extraction of real estate inventory data of the present invention;
[0037] Figure 3 is the interface diagram of real estate risk early warning of the present invention;
[0038] Figure 4 is the schematic diagram of the principle of the topology diagram of the present invention;
[0039] Figure 5 is the structural schematic diagram of the real estate data processing and early warning system based on big data of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0040] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0041] It can be understood that the terms "first", "second", etc. used in the present application can be used in this article to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish the first element from another element. For example, without departing from the scope of the present application, the first xx script can be called the second xx script, and similarly, the second xx script can be called the first xx script.
[0042] As Figure 1 shown, a real estate data processing and early warning method based on big data includes:
[0043] S1: Determine the first data source and the second data source, perform an overall extraction of the stock data of the first data source and the second data source to obtain the first main data and the first additional data.
[0044] Specifically, the first data source is the e-government network, and the second data source is third-party Internet resources. First, use Kettle to download the real estate data of all current enterprises from the e-government network to obtain the first main data, and download the extended data related to the enterprises from the third-party Internet resources as the first additional data. The first additional data includes, for example, the enterprise real estate information provided by third-party websites, as well as the business conditions, traffic conditions, building age, property information, etc. around the real estate information.
[0045] S2: Write the first main data and the first additional data into the database, perform data governance on the first main data in combination with the first additional data to obtain the second main data, and perform data reduction on the first additional data after the governance is completed to obtain the second additional data.
[0046] The data governance in this embodiment includes duplicate data deletion, missing data filling, data transcoding, and data desensitization.
[0047] The following takes the first data source as an example for detailed description. As Figure 2 shown, during the data writing process, first write the data into the ODS library (operational data store) of MYSQL. Specifically, use DS to call the kettle task for real estate stock data to extract and store the data. After DS successfully calls kettle to extract the real estate stock data, as Figure 2 shown, Kettle is an open-source ETL tool. Through the configuration of the graphical interface, data migration and data governance can be achieved without developing code. DS is short for DolphinSchedule, which is a distributed and easily extensible visual workflow task scheduling platform. After the data extraction is completed, data governance is performed on the first main data. The data quality mainly includes removing duplicate data, filling in null data and missing data, setting special marks for data that cannot be supplemented, performing parameter conversion for variable data or code-type data according to the code value table, and desensitizing sensitive data. Sensitive data includes, for example, enterprise names, unified social credit codes, etc. In addition, when filling in data, search and supplement in the first additional data according to the known real estate information. After using the first additional data to complete the filling of the first main data, it is necessary to perform data reduction on the first additional data to avoid occupying a large storage space.
[0048] S3: Set the first attribute, extract the first attribute from the second main data to generate a first form for each enterprise, extract the second attribute from the first form to generate a second form, the second attribute includes the real estate mortgage status, the real estate seizure status, the real estate valuation value and the number of mortgages.
[0049] In this embodiment, the first attribute specifically includes the name of the right holder, the right holder's certificate type, the right holder's certificate number, the real estate certificate number, the real estate unit number, the location, the region, the nature of the rights, the purpose, the building area, the registration date (issuance date), the measured floor area, the measured apportioned area, the nature of the house, the real estate seizure status, the seizure period, the name of the executor, the name of the person to be executed, the real estate mortgage status, the mortgage amount, the mortgage period, the number of mortgages, the name of the house owner (mortgagor), the house owner (mortgagor) certificate type, the house owner (mortgagor) certificate number, the real estate valuation, the update time and other parameters; then the second attribute is extracted from the above-mentioned first attribute to generate a second form, and the second attribute includes the real estate mortgage status, the real estate seizure status, the real estate valuation value and the number of mortgages.
[0050] S4: extracting the first main data and the first additional data again at every preset time interval to update the second form in the database.
[0051] Before updating, first set the data update task and the corresponding execution time. Specifically, use DS to schedule Kettle's update task on a daily basis and set up error email reminders. After building the extraction and update mechanism for each data source, apply kettle's "Insert Update Component" to perform a full-field comparison of the source data with the ODS stock data, thereby inserting the new data into ODS and completing the update of the second form.
[0052] S5: Setting a judgment condition, when the second attribute in the second form satisfies the judgment condition, triggering a risk warning.
[0053] In this embodiment, when the mortgage status of the real estate changes, the seizure status of the real estate changes, the change value of the real estate valuation exceeds the fourth threshold, and the number of mortgages increases, it is determined that the second attribute meets the judgment condition.
[0054] When the second attribute shows that the mortgage status of the real estate has changed, the seizure status of the real estate has changed, the change in the real estate valuation value exceeds the fourth threshold, or the number of mortgages increases, it is judged that the second attribute meets the judgment condition. For example, the fourth threshold is set to 10000, and the following will be generated: Figure 3 The risk warning diagram shown is used to remind relevant personnel to pay attention to enterprise risks.
[0055] The present invention obtains the real estate data of enterprises in various ways, and selects the data downloaded from the official website as the main data, that is, the first subject data. Before analysis, data governance is carried out on the first subject data to ensure its data quality. During the governance process, if there are data missing in the main data, the first additional data other than the first subject data is used to supplement the missing values, thus ensuring the integrity of the first subject data. The present invention enables relevant personnel to quickly obtain the real estate information of enterprises in multiple dimensions, and realizes the real-time tracking and early warning of abnormal real estate information of enterprises by setting judgment conditions.
[0056] After the present invention supplements the missing values of the first subject data with the first additional data, it also streamlines the first additional data, thereby reducing the occupied space of the first additional data, accelerating the running speed of the database, and reducing the expenditure cost of the hardware.
[0057] It should be particularly noted that through the present invention, the real estate situation of enterprises can be analyzed, making up for the problem of incomplete analysis in the existing technology.
[0058] The data streamlining of the first additional data by the present invention includes the following steps:
[0059] The second data source includes multiple sub-data sources. Locate the target data of numerical type in the sub-data sources, organize the target data belonging to the same data name in the sub-data sources into a numerical sequence, calculate the first similarity between different data names in different sub-data sources, and divide the data names into multiple first data groups. The first similarity between any two data names in the first data group is greater than the first threshold.
[0060] The sub-data sources are different third-party websites. After grabbing data from different third-party websites, locate the target data of numerical type therein, such as real estate valuation, real estate transaction price, reference average price per square meter, etc. Then, various numerical values belonging to the same name will be organized into a numerical sequence about time. For example, the daily real estate valuations on website A are organized into a numerical sequence.
[0061] The purpose of this step is to merge the same type of data from different third-party websites, such as real estate valuation data. Specifically, if the data name related to real estate valuation on website A is "estimated total price of the house", and the data name related to real estate valuation on website B is "estimated selling price of the house", then compare the first similarity between "estimated total price of the house" and "estimated selling price of the house". In this example, the number of identical characters in the two names is used as the first similarity. For example, in the above two names, each name has 6 characters, and 5 of them are the same, so the first similarity is 5 / 6 = 0.83, and the first threshold is 0.5. Here, the first similarity is greater than the first threshold.
[0062] Calculate the second similarity between different numerical sequences in the first data group, construct a topological graph based on the second similarity. The topological graph includes multiple nodes, where the nodes represent the numerical sequences in the first data group, and the optimal similarity is marked between the nodes.
[0063] Due to the existence of multiple sub-data sources, there may be more than 2 numerical sequences in each first data group. For example, for real estate valuation, the numerical sequences from websites A, B, and C are A, B, and C respectively, and they are all in one first data group. Then the first similarities between numerical sequences A and B, B and C, and A and C are all greater than the first threshold. The difference between the corresponding data at the same time point in two numerical sequences can be calculated, and the corresponding variance can be calculated based on all the calculated differences, and the variance is used as the second similarity between the two numerical sequences. Through this step, sequences with similar names and similar data can be grouped into one data group. The constructed topological graph is as Figure 4 shown. The construction method of the topological graph will be introduced later. The nodes A, B, and C in the topological graph correspond to the numerical sequences A, B, and C respectively.
[0064] Screen two adjacent nodes in the topological graph, determine the first sequence and the second sequence in the numerical sequences corresponding to the two nodes, and generate a mapping function based on the first sequence and the second sequence. The mapping function is used to transform the first sequence into the second sequence.
[0065] The database deletes the second sequence in the sub-data source, retains the first sequence corresponding to the deleted second sequence and the mapping function to complete the reduction of the first additional data.
[0066] Determine which of the two numerical sequences corresponding to nodes A and B is the first sequence and which is the second sequence, and then generate a mapping function according to the first sequence and the second sequence. Here, the mapping function is a linear function, such as y = ax + b, where a ∈ E, E is the set of polynomial coefficients of the mapping function, b is a constant, x is the first sequence, and y is the second sequence. Existing linear regression tools can be used to determine E and b, so as to determine the mapping function of the first sequence and the second sequence. Specifically, when converting, the values in the first sequence are input into the mapping function to obtain the second sequence.
[0067] Through the previous screening, two sequences with extremely similar numerical distributions can be obtained. Especially in this embodiment, the variance is used as the second similarity, making the distribution of the two numerical sequences closer after screening. Even if the restored second sequence is different from the original second sequence, the difference will be very small. On this basis, the data is reduced by deleting the second sequence. When subsequent other aspects of data analysis need to use the first additional data, it can be restored with the help of the first sequence and the mapping function without having to obtain it from the website again.
[0068] The steps for constructing the topology graph in this embodiment include the following:
[0069] Extract two numerical sequences corresponding to the maximum second similarity from the first data group, define them as the first initial sequence and the second initial sequence, and define the corresponding nodes as the first initial point and the second initial point respectively. Screen the numerical sequence with the maximum second similarity to the first initial sequence or the second initial sequence from the first data group as the third initial sequence, and connect it as the third initial point to the corresponding first initial point or second initial point. Repeat this step until all numerical sequences in the first data group are extracted to obtain the topology graph corresponding to the first data group.
[0070] For example, the first data group includes numerical sequences A, B, C, and D. There will be 6 second similarities among them. Obtain the second similarity with the largest value, such as the second similarity between numerical sequence A and numerical sequence C is the largest, and define them as the first initial point and the second initial point respectively; then obtain the second similarity between numerical sequences B and D and numerical sequence A, and the second similarity between numerical sequences B and D and numerical sequence C. After comparison, the second similarity between numerical sequence B and A is greater than the second similarity between numerical sequence B and C, and both are greater than the second similarity between numerical sequence D and A or C. Then, take numerical sequence B as the third initial sequence, and in the topology graph, connect the third initial sequence as the third initial point to the first initial point. Finally, taking numerical sequence C as the target, analyze the numerical sequence among A, B, and D with the largest second similarity to it. For the topology graph obtained by this method, the second similarity between adjacent nodes is the largest, aiming to reduce the error value of restoring the second sequence from the first sequence later and further improve the accuracy of data restoration.
[0071] The steps for generating the mapping function in this embodiment include the following:
[0072] Define the numerical sequences corresponding to two adjacent nodes in the topology graph as the first alternative sequence and the second alternative sequence. Generate the first function and the second function based on the first alternative sequence and the second alternative sequence. Use the first function to transform the first alternative sequence into the third alternative sequence, and use the second function to transform the second alternative sequence into the fourth alternative sequence. Calculate the first difference degree between the first alternative sequence and the fourth alternative sequence, and the second difference degree between the second alternative sequence and the third alternative sequence. If the first difference degree is less than the second difference degree, define the first alternative sequence as the first sequence and the first function as the mapping function; otherwise, define the second alternative sequence as the first sequence and the second function as the mapping function.
[0073] Continue to refer to Figure 4, in this embodiment, the first sequence, the second sequence, and the mapping function are determined through the following steps: First, select node A and node B, corresponding to the first alternative sequence and the second alternative sequence respectively, and generate a first function for converting the first alternative sequence into the second alternative sequence and a second function for converting the second alternative sequence into the first alternative sequence. Specifically, in linear regression analysis, the mutual exchange of the dependent variable and the independent variable will affect the subsequently generated mapping function, and the change of the mapping function will affect the reduction accuracy. Based on this, after generating the first function and the second function, use the first function to convert the first alternative sequence into a third alternative sequence, calculate the first difference degree between the third alternative sequence and the second alternative sequence, use the second function to convert the second alternative sequence into a fourth alternative sequence, and calculate the second difference degree between the fourth alternative sequence and the first alternative sequence. The sum of the absolute values of the differences in the corresponding positions of the two numerical sequences can be used as the difference value. For example, if the first alternative sequence is 1, 2, 3 and the third alternative sequence is 2, 1, 3, then the difference degree is |1 - 2| + |2 - 1| + |3 - 3| = 2. Finally, select the function corresponding to the smaller difference degree as the mapping function, the alternative sequence used for conversion is the first sequence, and the converted alternative sequence is the second sequence. Then select node B and node C again for analysis to determine the first sequence and the second sequence among them.
[0074] After this embodiment simplifies the topological graph, if there are a first reserved point and a second reserved point, all nodes between the first reserved point and the second reserved point are deleted, and the deleted nodes are defined as target points. Starting from the first reserved point and the second reserved point respectively, combine the optimal similarity to calculate the cumulative product value of the reserved points. If the cumulative product values of all reserved points are less than the third threshold, then restore the reserved point.
[0075] Continue to refer to Figure 4 , during deletion, among nodes A and B, if A is the first sequence, then the second sequence corresponding to node B will be deleted. Among nodes B and C, node B is the first sequence, and the second sequence corresponding to node C will be deleted. Among nodes C and D, node C is the first sequence, and the second sequence corresponding to node D will be deleted. Among nodes D and E, node E is the first sequence, and the second sequence corresponding to node D will be deleted. Eventually, nodes B, C, and D will be deleted. Correspondingly, nodes A and node B will become the first reserved point and the second reserved point, and nodes B, C, and D will be located as target points.
[0076] If the optimal similarity between node A and node B is 0.92, the optimal similarity between node B and node C is 0.8, the optimal similarity between node C and node D is 0.85, the optimal similarity between node D and node E, and the optimal similarity between node A and node B is 0.92. For node C, the cumulative product value from node A to node C is 0.92 0.8 = 0.736, and the cumulative product value from node E to node C is 0.92 0.85 = 0.782. Assuming the third threshold is 0.75, then node C is restored.
[0077] Since restoring node A to node B will result in certain errors, and restoring from node B to node C will amplify the existing errors, which will lead to a decreasing accuracy in subsequent restorations. Therefore, the purpose of this step is to avoid an overly long restoration chain that continuously amplifies the error value.
[0078] The missing value filling for the first main data in this embodiment includes the following steps:
[0079] Obtain the first main data with missing values, obtain the key information of the first main data, and search in the first additional data based on the key information for missing value filling.
[0080] Specifically, for example, if the real estate valuation value in the first main data is missing, then obtain the key information such as the location, belonging area, and name of the right holder of the real estate, and use this key information to retrieve in the first additional data. If a real estate valuation value that matches the key information is obtained in the first additional data, then import it into the first main data to complete the missing value filling.
[0081] In this embodiment, the first additional data redrawn is defined as new data. Whenever the number of new data reaches the fifth threshold, data reduction is performed on the currently existing new data.
[0082] As Figure 5 shown, the present invention also provides a real estate data processing and early warning system based on big data. This system is used to implement the above-mentioned real estate data processing and early warning method based on big data. The system includes:
[0083] A collection module that determines the first data source and the second data source. The collection module is used to perform an overall extraction of the stock data of the first data source and the second data source to obtain the first main data and the first additional data, and extract the first main data and the first additional data again at preset time intervals to update the second form in the database;
[0084] A storage module that writes the first main data and the first additional data into the database, performs data governance on the first main data in combination with the first additional data to obtain the second main data, and performs data reduction on the first additional data after the governance is completed to obtain the second additional data;
[0085] An extraction module sets a first attribute, extracts the first attribute from second body data to generate a first form for each enterprise, and extracts a second attribute from the first form to generate a second form. The second attribute includes real estate mortgage status, real estate seizure status, real estate valuation value, and number of mortgages;
[0086] An early warning module sets a judgment condition. When there is a second attribute in the second form that meets the judgment condition, a risk warning is triggered.
[0087] It should be understood that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0088] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for real estate data processing and early warning based on big data, characterized in that Including: Determine the first data source and the second data source, and perform an overall extraction of the stock data of the first data source and the second data source to obtain first main data and first additional data; Write the first main data and the first additional data into a database, perform data governance on the first main data in combination with the first additional data to obtain second main data, and perform data reduction on the first additional data after the governance is completed to obtain second additional data; Set a first attribute, extract the first attribute from the second main data to generate a first form for each enterprise, and extract a second attribute from the first form to generate a second form, where the second attribute includes real estate mortgage status, real estate seizure status, real estate valuation value, and number of mortgages; Extract the first main data and the first additional data again at preset time intervals to update the second form in the database; Set a judgment condition, and trigger a risk warning when the second attribute in the second form meets the judgment condition; The data reduction of the first additional data includes the following steps: The second data source includes multiple sub-data sources. Locate the target data of numerical type in the sub-data sources, organize the target data belonging to the same data name in the sub-data sources into a numerical sequence, calculate the first similarity between different data names in different sub-data sources, and divide the data names into multiple first data groups, where the first similarity between any two data names in the first data group is greater than a first threshold; Calculate the second similarity between different numerical sequences in the first data group, and construct a topological graph based on the second similarity. The topological graph includes multiple nodes, where the nodes represent the numerical sequences in the first data group, and the optimal similarity is marked between the nodes; Screen two adjacent nodes in the topological graph, determine a first sequence and a second sequence in the numerical sequences corresponding to the two nodes, and generate a mapping function based on the first sequence and the second sequence. The mapping function is used to transform the first sequence into the second sequence; The database deletes the second sequence in the sub-data source, retains the first sequence corresponding to the deleted second sequence and the mapping function to complete the reduction of the first additional data.
2. The method according to claim 1, wherein Constructing the topological graph includes the following steps: Extract the two numerical sequences corresponding to the maximum second similarity from the first data group, define them as a first initial sequence and a second initial sequence, and define the corresponding nodes as a first initial point and the second initial point respectively. Screen the numerical sequence with the maximum second similarity to the first initial sequence or the second initial sequence from the first data group as a third initial sequence, and connect it as a third initial point to the corresponding first initial point or second initial point. Repeat this step until all the numerical sequences in the first data group are extracted to obtain the topological graph corresponding to the first data group.
3. The method according to claim 1, characterized in that, Generating the mapping function includes the following steps: Two numerical sequences corresponding to two adjacent nodes in the topological graph are defined as a first alternative sequence and a second alternative sequence. A first function and a second function are generated based on the first alternative sequence and the second alternative sequence. The first alternative sequence is transformed into a third alternative sequence using the first function, and the second alternative sequence is transformed into a fourth alternative sequence using the second function. A first difference degree between the first alternative sequence and the fourth alternative sequence, and a second difference degree between the second alternative sequence and the third alternative sequence are calculated. If the first difference degree is less than the second difference degree, the first alternative sequence is defined as the first sequence, and the first function is defined as the mapping function; otherwise, the second alternative sequence is defined as the first sequence, and the second function is defined as the mapping function.
4. The method according to claim 3, wherein After streamlining the topological graph, if there are a first reserved point and a second reserved point, and all nodes between the first reserved point and the second reserved point are deleted, the deleted nodes are defined as target points. Starting from the first reserved point and the second reserved point respectively, the cumulative product value of the reserved points is calculated in combination with the optimal similarity. If the cumulative product values of the reserved points are all less than a third threshold, then the reserved points are restored.
5. The method according to claim 1, wherein Missing value filling for the first main data includes the following steps: Obtain the first main data with missing values, obtain the key information of the first main data, and search in the first additional data based on the key information to fill the missing values.
6. The method according to claim 5, wherein The data governance includes duplicate data deletion, missing data filling, data transcoding, and data desensitization.
7. The method according to claim 1, characterized in that When the real estate mortgage status changes, the real estate seizure status changes, the change value of the real estate valuation numerical value exceeds a fourth threshold, or the number of mortgages increases, it is determined that the second attribute meets the judgment condition.
8. The method according to claim 1, characterized in that, The first additional data re-extracted is defined as new data. Whenever the number of new data reaches a fifth threshold, the existing new data is streamlined.
9. A real estate data processing and early warning system based on big data, which is used to implement the method described in any one of claims 1-8, and is characterized in that, Including: A collection module determines a first data source and a second data source. The collection module is used to extract the stock data of the first data source and the second data source as a whole to obtain first main data and first additional data, and re-extract the first main data and the first additional data at preset time intervals to update the second form in the database. A storage module writes the first body data and the first additional data into the database, performs data governance on the first body data in combination with the first additional data to obtain second body data. Wherein, the second data source includes multiple sub-data sources, locates target data of numerical type in the sub-data sources, arranges the target data belonging to the same data name in the sub-data sources into a numerical sequence, calculates a first similarity between different data names in different sub-data sources, divides the data names into multiple first data groups, the first similarity between any two data names in the first data group is greater than a first threshold, calculates a second similarity between different numerical sequences in the first data group, constructs a topological graph based on the second similarity, the topological graph includes multiple nodes, the nodes represent the numerical sequences in the first data group, and the optimal similarity is marked between the nodes, screens two adjacent nodes in the topological graph, determines a first sequence and a second sequence in the numerical sequences corresponding to the two nodes, generates a mapping function based on the first sequence and the second sequence, the mapping function is used to convert the first sequence into the second sequence, the database deletes the second sequence in the sub-data source, retains the first sequence corresponding to the deleted second sequence and the mapping function to complete the data streamlining governance of the first additional data, and performs data streamlining on the first additional data after completion to obtain second additional data; An extraction module sets a first attribute, extracts the first attribute from the second body data to generate a first form for each enterprise, and extracts a second attribute from the first form to generate the second form, the second attribute includes real estate mortgage status, real estate seizure status, real estate valuation value and number of mortgages; An early warning module is set with a judgment condition, and when there is the second attribute in the second form that meets the judgment condition, a risk warning is triggered.
Citation Information
Patent Citations
Enterprise risk analysis method and system based on big data, and medium
CN116468271A
Supplier enterprise risk assessment method and system based on big data technology
CN118153964A
Method and system for determining information similarity
CN112989211A
Data processing method and device
CN113032399A