An intelligent tax assistant interaction method
Through laser scanning and panoramic cameras, the internal space layout diagram of the building is obtained, combined with machine learning and cluster analysis technology, the tax classification problem of mixed use in the building is solved, and the rapid and accurate tax calculation and analysis report generation is achieved, which improves the intelligence and accuracy of tax management.
Patent Information
- Application Number
- CN202510175587.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-02-18
AI Technical Summary
After the internal partitions of civil real estate are renovated, there may be mixed use of office and residential areas within the same building, resulting in challenges in tax classification data processing. Traditional manual measurement and calculation methods are time-consuming and labor-intensive and difficult to ensure the accuracy and consistency of data.
The internal spatial layout diagram and area data of the building are obtained through laser scanning and panoramic cameras, and the geometric features and spatial relationships of each area are extracted in combination with image segmentation and object detection technology, and the use properties are judged, and the tax classification and tax calculation model is established using supervised machine learning, cluster analysis, case reasoning and logistic regression algorithms.
It realizes the rapid and accurate identification and division of space areas for different purposes within the building, calculates the tax amount to be paid according to the corresponding tax rate standards, and generates a detailed tax analysis report, which improves the intelligence and precision of tax management.
Smart Images

Figure CN119648448B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to an intelligent tax assistant interaction method. Background Art
[0002] After the internal partition of civil buildings is renovated, there may be a situation where offices and residences are mixed within the same building, which poses significant data processing challenges for tax classification. When an intelligent tax assistant divides the areas of different uses and calculates the tax amount, it faces the technical problem of how to accurately define the office use and residential use. When the same building becomes a mixed-use property after being partitioned, how to accurately divide the areas of different uses and apply the corresponding tax rates respectively becomes a thorny issue. The mixed use of offices and residences may make it difficult to separately measure the water and electricity costs, further increasing the complexity of tax amount calculation. Traditional manual measurement and calculation methods are not only time-consuming and laborious, but also difficult to ensure the accuracy and consistency of data. If a unified tax rate is simply determined according to the main use of the building, it may lead to tax unfairness and is difficult to truly reflect the actual use situation of the building. Therefore, there is an urgent need for an intelligent technical means that can quickly and accurately identify and divide the spatial areas of different uses inside the building, calculate the payable tax amount according to the corresponding tax rate standards, and generate a detailed tax analysis report to provide a reliable basis for the tax collection and management work of the tax department. Summary of the Invention
[0003] The present invention provides an intelligent tax assistant interaction method, mainly including:
[0004] Through laser scanning and panoramic cameras, obtain the spatial layout diagram and area data of each region after the internal partition renovation of the building, perform image segmentation and object detection on the spatial layout diagram, extract the geometric features and spatial relationships of each region, and combine with building design specifications and tax classification standards to judge the usage nature of each region. The usage nature includes office, residential, and mixed use;
[0005] If the intelligent tax assistant determines that the usage nature of the region is office or residential, then use a supervised machine learning algorithm to train the existing tax classification sample data to establish a tax classification model for judging the tax rate categories corresponding to the office area and the residential area;
[0006] If the intelligent tax assistant determines that the usage nature of the region is a mixed nature, then through a clustering analysis algorithm, divide the mixed nature region into several sub-regions, and determine the sub-regions as office nature or residential nature according to the area ratio and usage characteristics of the sub-regions;
[0007] Through the case-based reasoning algorithm, retrieve similar cases from the existing tax case database for the sub-region situation, compare and analyze relevant attributes, including regional area, usage function, and personnel occupancy, predict the tax rate tendency of the region, and form preliminary tax suggestions;
[0008] Through the logistic regression algorithm, analyze the influence weights of the area and tax rate of each region on the comprehensive tax rate, calculate the comprehensive tax rates of the office area and residential area using the weighted average method, and determine the proportion coefficients of each region in the tax amount calculation;
[0009] Integrate the area, tax rate, and proportion coefficients of each region, use the dynamic programming algorithm to establish a tax amount calculation model, obtain the payable tax amount of the building, and at the same time evaluate the influence degree of the changes in the regional area and tax rate on the tax amount through the sensitivity analysis method, and form a tax optimization plan through the intelligent tax assistant;
[0010] Generate an intelligent tax suggestion report in the intelligent tax assistant according to the comprehensive tax rate, payable tax amount, and tax optimization plan of the building.
[0011] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0012] The present invention discloses an intelligent tax assistant interaction method for tax classification and tax amount calculation of the space after the internal renovation of a building. Obtain the space layout diagram through laser scanning and panoramic cameras, extract the features of each region using image segmentation and object detection, and judge the usage nature in combination with building codes. For office and residential areas, establish a tax classification model using supervised learning; for mixed-property areas, subdivide the sub-regions through cluster analysis. Also use case-based reasoning to retrieve similar cases from the case library to predict the tax rate tendency; analyze the influence weights of various factors on the tax rate through logistic regression, and establish a tax amount calculation model using dynamic programming. Finally, generate an intelligent tax suggestion report to provide decision support for the tax classification and tax amount calculation of the building, and improve the intelligent and precise level of tax management. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 It is a flowchart of an intelligent tax assistant interaction method of the present invention.
[0014] Figure 2 It is a schematic diagram of an intelligent tax assistant interaction method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0016] As Figure 1-2 , a specific intelligent tax assistant interaction method in this embodiment may include:
[0017] In step S101, a spatial layout map of the interior partition renovation of the building and area data of each region are obtained through laser scanning and a panoramic camera. Image segmentation and object detection are performed on the spatial layout map to extract the geometric features and spatial relationships of each region. Combining the building design specifications and the tax classification standards, the usage nature of each region is determined, and the usage nature includes office, residential, and mixed use.
[0018] The first point cloud data inside the building is obtained by using a three-dimensional laser scanner, and a first texture point cloud model is obtained through registration processing of the first point cloud data and panoramic image data; spatial clustering processing is performed on the first texture point cloud model to obtain a second texture point cloud model, and based on the second texture point cloud model, the wall surfaces are extracted by using the region growing algorithm and the relative position relationship between the walls is calculated to obtain a first spatial connectivity graph; vectorization processing is performed on the first spatial connectivity graph to obtain a first floor plan, and image segmentation processing is performed on the first floor plan to obtain a second floor plan; shape feature vectors and spatial topology feature vectors are extracted from the second floor plan, and a room usage classification model is trained by using the support vector machine algorithm for the shape feature vectors and spatial topology feature vectors, and the office and residential usage of the rooms are labeled according to the classification model.
[0019] Specifically, the interior of the building is scanned from multiple acquisition points by a 3D laser scanner to collect the first point cloud data. The panoramic camera synchronously collects panoramic image data according to the acquisition point positions. The first point cloud data is subjected to registration and splicing processing to obtain the second point cloud data. The projection transformation algorithm is used to perform geometric correction on the panoramic image data and register it with the second point cloud data to obtain the first texture point cloud model. The point cloud data is denoised and spatially clustered according to the first texture point cloud model to obtain the second texture point cloud model. The region growing algorithm is used to extract the interior wall surfaces and floor surfaces, and the least squares plane fitting method is used to obtain the wall thickness parameters and floor elevation parameters. The relative position relationship between the walls is calculated through spatial topology analysis to obtain the first spatial connectivity graph. Based on the first spatial connectivity graph, vectorization processing is used to obtain the first floor plan. Image segmentation is performed on the first floor plan to obtain the second floor plan. The wall contour lines are extracted from the second floor plan and the positions of the door and window openings are identified. The polygon area calculation method is used to calculate the room usable area data. The room shape feature vectors are extracted according to the second floor plan, including the room length-width ratio, area ratio, and perimeter ratio. The spatial topology feature vectors are extracted based on the first spatial connectivity graph, including the number of adjacent rooms, door and window connectivity, and corridor width. The support vector machine algorithm is used to establish a room use classification model. The classifier is trained through the shape feature vectors and topology feature vectors. The classification threshold is set according to the minimum usable area of the office space and the residential design specifications, and the rooms are labeled with office, residential, and mixed use properties. A second spatial connectivity graph is constructed according to the labeling results. The graph theory algorithm is used to analyze the spatial organization relationship of rooms with different uses, calculate the aggregation degree of rooms with the same use, and judge the dominant use properties of each area of the building. In the field of determining the use properties after the interior renovation of the building, when collecting point cloud data by a 3D laser scanner, the horizontal angular resolution of the scanner is set to 0.02 degrees, the vertical angular resolution is set to 0.02 degrees, the scanning distance is set to 0.5 m to 30 m, multiple scanning stations are set inside the building, and the distance between adjacent stations is 3 m to 5 m to ensure that the scanning data overlap rate is above 40%. The resolution of the images collected by the panoramic camera is 8000×4000 pixels, the horizontal field of view angle of a single image is 360 degrees, and the vertical field of view angle is 180 degrees to ensure that the images have no distortion and color difference. The iterative closest point algorithm is used for point cloud data registration. The registration error between two sets of point cloud data is calculated in each iteration step, and the iteration stops when the error is less than 1 mm to obtain the accurately registered point cloud data. When registering the panoramic image and the point cloud data, the image feature points and point cloud feature points are extracted for matching, and a projection transformation matrix is established to achieve the accurate mapping of the image texture to the point cloud model. When extracting the wall surfaces and floor surfaces, the region growing algorithm sets the normal vector angle threshold to 3 degrees and the plane fitting residual threshold to 2 mm to reliably extract the main plane components inside the building. The least squares method is used to fit the wall plane equation, ax + by + cz + d = 0, where (a, b, c) is the plane normal vector and d is the distance from the plane to the origin.The wall thickness is obtained by calculating the distance between adjacent wall planes, and the typical wall thickness ranges from 120 mm to 240 mm. When extracting the shape characteristics of a room, the aspect ratio of the room is calculated as the maximum side length / the minimum side length, the area ratio is the usable area / the building area, and the perimeter ratio is the actual perimeter / the perimeter of the minimum circumscribed rectangle. Taking an office area as an example, the aspect ratio of the room is 1.8, the area ratio is 0.85, and the perimeter ratio is 1.2, indicating that the room is a regular rectangle. In the spatial topological characteristics, the number of adjacent rooms is 4, the door and window connectivity is 6, and the corridor width is 1500 mm, meeting the design requirements of the office space. The support vector machine classifier uses the radial basis kernel function, K(x,y)=exp(-γ||x - y||²), where γ is the kernel function parameter and K(x,y) is the kernel function value, representing the similarity between two input vectors x and y in the feature space. The classification threshold is set such that the usable area of office space is not less than 4 square meters per person, and the usable area of a residential bedroom is not less than 9 square meters. In a certain building, 42% is designated as the office area, 35% as the residential area, and 23% as the mixed-use area. Through spatial aggregation analysis, it is determined that the dominant use nature of this building is office.
[0020] Step S102, if the intelligent tax assistant determines that the use nature of the area is office or residential, a supervised machine learning algorithm is used to train the existing tax classification sample data to establish a tax classification model for determining the tax rate categories corresponding to the office area and the residential area.
[0021] The standard deviation method is used to identify outliers in the tax classification sample data. According to the outlier identification results, the mean filling method is used to supplement the missing values to obtain the first sample data; the minimum-maximum normalization process is performed on the first sample data, and the one-hot encoding transformation is performed on the normalization result to obtain the third sample data; for the third sample data, the variance selection method is used to remove the low-discrimination features, and the principal components with a cumulative contribution rate reaching the specified threshold are extracted from the data after removing the low-discrimination features through the principal component analysis method to obtain the sixth sample data; a random forest classifier is established based on the sixth sample data, and the tax rate category probability value is calculated through the random forest classifier. If the probability value is greater than the specified threshold, the corresponding tax rate category is adopted; if the probability value is less than the specified threshold, the tax rate category is determined by the probability weighting method.
[0022] Specifically, outlier processing is performed on the tax classification sample data through data cleaning. The three - standard - deviation method is used to identify outliers, and the mean filling method is used to supplement missing values to obtain the first sample data. The minimum - maximum normalization processing is performed on the numerical features in the first sample data to obtain the second sample data, and one - hot encoding transformation is performed on the categorical features to obtain the third sample data. The Pearson correlation coefficient is calculated based on the third sample data, and redundant features with a correlation coefficient greater than 0.9 are removed to obtain the fourth sample data. The variance selection method is used to remove low - discriminant features with a variance less than 0.1 to obtain the fifth sample data. Key features such as building area, service life, and regional value are extracted for the fifth sample data. The principal component analysis method is used to calculate the feature contribution rate, and the principal components with an accumulated contribution rate reaching 90% are selected to obtain the sixth sample data. A random forest classifier is established based on the sixth sample data, with the number of decision trees set to 100, the maximum depth of the tree set to 10, and the minimum number of samples in the leaf nodes set to 5. The five - fold cross - validation method is used to train the classifier to obtain the first tax classification model. The new area is predicted according to the first tax classification model, and the probability values of each tax rate category are calculated. If the maximum probability value is greater than 0.8, the corresponding tax rate category is adopted; if the maximum probability value is less than 0.8, the tax rate category is determined by probability weighting. For the mixed - use area, the proportion of office - use area and the proportion of residential - use area are calculated respectively, and the area ratio is used to weighted - average the tax rates corresponding to different uses to obtain the final tax rate. During the processing of the tax classification sample data, the building area data of the property shows an obvious outlier distribution. By calculation, the mean of the building area is 120 square meters, and the standard deviation is 30 square meters. The three - standard - deviation method is used to identify records with a building area greater than 210 square meters or less than 30 square meters as outliers. There are 15% missing values in the regional value feature. By calculating the mean of this feature as 12,000 yuan per square meter, this mean is used to fill the missing data. Correlation analysis is performed on 11 original features such as the service life of the property, housing structure type, and decoration grade. By calculating the Pearson correlation coefficient, it is found that the correlation coefficient between the service life and the housing condition rate is 0.92, indicating that these two features have a strong linear correlation, and the service life feature is retained. The housing structure type includes reinforced concrete, brick - concrete, and brick - wood, and is transformed into three binary features by one - hot encoding. During the feature selection process, the variance values of each feature are calculated, and it is found that the variance of the housing orientation feature is 0.08, which is less than the threshold of 0.1, indicating that this feature has a low discriminant ability for the samples. The principal component analysis method is used to reduce the dimension of the remaining features. The feature contribution rate of the first principal component is 45%, the second principal component is 28%, and the third principal component is 17%. The accumulated contribution rate reaches 90%, and these three principal components are selected as classification features. The training process of the random forest classifier uses five - fold cross - validation, that is, the sample data is randomly divided into 5 parts, and 4 parts are selected as the training set each time, and 1 part is used as the validation set.Training is carried out by setting different parameter combinations. When the number of decision trees is 100, the maximum depth is 10, and the minimum number of samples in a leaf node is 5, the classification accuracy reaches the optimal value of 92%. For a newly added area with a building area of 150 square meters, the probability of the office building tax rate category is 0.85, and the probability of the residential building tax rate category is 0.15. Since the maximum probability value of 0.85 is greater than the threshold of 0.8, it is determined that the area applies the office building tax rate. For another mixed-use area with a building area of 200 square meters, where the office area is 120 square meters and the residential area is 80 square meters, the proportion of the office area is 0.6, and the proportion of the residential area is 0.4. The office building tax rate is 12%, and the residential building tax rate is 4%. The comprehensive tax rate of this area is calculated to be 8.8% through weighted average.
[0023] Step S103, if the intelligent tax assistant determines that the usage nature of the area is a mixed nature, then through the clustering analysis algorithm, the mixed-nature area is subdivided into several sub-areas, and according to the area proportion and usage characteristics of the sub-areas, it is determined whether the sub-areas are of office nature or residential nature.
[0024] By collecting the basic feature data of the door and window positions, partition walls, and passage connections in the mixed-nature area, the first feature vector is obtained by using the maximum-minimum normalization method; the Euclidean distance is calculated based on the first feature vector to obtain the first similarity matrix, and the density clustering algorithm is used to obtain the first sub-area division result; for the first sub-area division result, the room connection relationship is extracted, and the second similarity matrix is obtained by calculating the passage width and connection method between adjacent rooms, and the graph cut algorithm is used to obtain the second sub-area division result; according to the second sub-area division result, the building area data is extracted, the crowd density data is collected by an infrared sensor, the hierarchical clustering algorithm is used to obtain the office nature clustering cluster and the residential nature clustering cluster, and the usage nature of the sub-area is determined according to the area weight and the clustering feature weight.
[0025] Specifically, for the mixed-use area, collect the first basic feature data such as the positions of doors and windows, partition walls, passage connections, water and electricity layouts, lighting and ventilation conditions, and equipment and facility configurations. Normalize the first basic feature data through the maximum-minimum normalization method to obtain the first feature vector. Calculate the Euclidean distance based on the first feature vector to obtain the first similarity matrix. Use the density clustering algorithm to subdivide the mixed-use area, set the clustering radius to 0.5, and the minimum number of sampling points to 4 to obtain the first sub-region division result. Based on the first sub-region division result, extract the room connectivity relationship, calculate the passage width and connection method between adjacent rooms to obtain the second similarity matrix, and use the graph cut algorithm to optimize the boundaries of the sub-regions that do not meet the connectivity requirements to obtain the second sub-region division result. For the second sub-region division result, extract the second basic feature data such as building area, partition type, and functional attributes, collect the pedestrian flow density data through infrared sensors, and extract the equipment usage frequency data based on the video surveillance data to obtain the second feature vector. Based on the second feature vector, use the hierarchical clustering algorithm to classify the sub-regions, set the number of clustering clusters to 2, and use the ward minimum variance method to obtain the office-use clustering cluster and the residential-use clustering cluster. Calculate the area data of each sub-region through the polygon area calculation method, calculate the proportion of the total area of the office-use area and the proportion of the total area of the residential-use area, and use an area weight of 0.7 and a clustering feature weight of 0.3 to determine the usage nature of the sub-regions. During the subdivision process of the mixed-use area, the basic feature data contains multiple dimensions: the positions of doors and windows are represented by coordinate points, such as (x1 = 350, y1 = 280) representing the upper left corner point of the door; the partition walls are represented by line segments, such as (x1 = 350, y1 = 280, x2 = 350, y2 = 580) representing a wall; the passage connections are recorded using an adjacency matrix, with a value of 1 indicating direct connection and a value of 0 indicating non-connection; the water and electricity layouts record the spatial distribution density of water supply points and power sockets. Normalize these raw data using the maximum-minimum normalization method to map all feature values to the [0, 1] interval. Taking a certain mixed area as an example, the door and window position coordinates are normalized from (350, 280) to (0.58, 0.47), the wall length is normalized from 300 cm to 0.6, and the passage width is normalized from 150 cm to 0.5. In density clustering, setting the clustering radius to 0.5 means that points with an Euclidean distance less than 0.5 are regarded as neighborhood points, and the minimum number of sampling points being 4 means that at least 4 neighborhood points are required to form a clustering core. The first sub-region division result obtained through calculation contains 5 sub-regions, where the largest sub-region area is 280 square meters and the smallest sub-region area is 85 square meters. The analysis of the room connectivity relationship shows that Sub-region 1 and Sub-region 2 are connected by a corridor with a width of 180 cm, Sub-region 2 and Sub-region 3 are connected by a door opening with a width of 90 cm, and Sub-region 4 and Sub-region 5 are connected by a passage with a width of 150 cm.Based on these connectivity features, the graph cut algorithm is used to optimize the boundary, cutting off the connection relationships with a channel width less than 120 cm, and re-dividing to obtain 4 sub-regions. In the extraction of functional attribute features, the infrared sensor is used to record the change of the pedestrian flow density within 24 hours. The pedestrian flow density in the office area is relatively high during the working hours from 9:00 to 18:00, and there are double peaks in the pedestrian flow density in the residential area in the morning and evening. The video surveillance data shows that the usage frequency of office equipment reaches the peak during working hours, while the usage frequency of living facilities is relatively high during non-working hours. When using hierarchical clustering, the distance between clusters is calculated by the ward minimum variance method, and the inter-class variance reaches the maximum value when the number of clustering clusters is 2. The total area of the office property area finally obtained is 320 square meters, accounting for 0.64; the total area of the residential property area is 180 square meters, accounting for 0.36. Combining the clustering eigenvalue, the comprehensive score is calculated according to the area weight of 0.7 and the feature weight of 0.3. The sub-region with a score greater than 0.55 is determined to be of office nature, and the one less than 0.45 is determined to be of residential nature.
[0026] Step S104, through the case-based reasoning algorithm, retrieve cases similar to the sub-region situation from the existing tax case database, compare and analyze relevant attributes, including regional area, usage function, and personnel occupancy, predict the tax rate tendency of this region, and form a preliminary tax recommendation.
[0027] Extract the first attribute data of building area, usage function, and traffic convenience from the tax case database, and obtain the feature vector through the min-max normalization method and the one-hot encoding method; use the cosine similarity method to retrieve cases with a similarity greater than the threshold from the tax case database according to the feature vector to obtain the first case set; extract the second attribute data of regional area ratio, space utilization rate, and personnel density for the first case set, and select the principal components with a cumulative contribution rate meeting the requirements through the principal component analysis method to obtain the second case set; based on the second case set, construct a decision tree model, use the information gain ratio to select the optimal splitting feature to extract the decision path, and obtain the tax rate prediction model through multi-layer perceptron training.
[0028] Specifically, extract the first attribute data such as building area, usage function, transportation convenience, decoration level, equipment and facilities, and personnel capacity from the tax case database. Normalize the numerical attributes through the min-max normalization method, and convert the categorical attributes through the one-hot encoding method to obtain feature vectors. Construct a distance metric matrix based on the feature vectors, and use the cosine similarity method to retrieve similar cases from the tax case database. Set the similarity threshold to 0.8, and select the cases with similarity greater than the threshold to form the first case set. Extract the second attribute data such as regional area ratio, space utilization rate, and personnel density for the first case set, and use the principal component analysis method to reduce the dimension of the second attribute data. Select the principal components with a cumulative contribution rate reaching 90% to obtain the second case set. Construct a decision tree based on the second case set, set the maximum depth of the tree to 5, and the minimum number of samples in the leaf nodes to 10. Use the information gain ratio to select the optimal splitting feature, and extract the decision path to obtain the first rule set. Calculate the rule confidence and support for the first rule set, set the confidence threshold to 0.7, and the support threshold to 0.1. Screen the rules that meet the threshold conditions to form the second rule set. Use a multi-layer perceptron to train the second rule set, set the number of hidden layer nodes to 10, the learning rate to 0.01, and the number of training rounds to 1000 to obtain the tax rate prediction model. Calculate the probability distribution of the sub-region under different tax rate categories based on the tax rate prediction model, and use the probability weighting method to obtain the tax rate prediction value. Combine the regional use attribute, area ratio attribute, and usage intensity attribute to correct the deviation of the prediction value. The original attribute data obtained from the tax case database contains multiple dimensions. The numerical range of the building area is between 50 and 1000 square meters, and it is converted to the interval from 0 to 1 through min-max normalization; the usage function includes three types: office, commercial, and residential, and is converted to three binary features through one-hot encoding; the transportation convenience is represented by the distance from the subway station, recorded as 1 within 300 meters, 0.6 from 300 to 800 meters, and 0.2 above 800 meters. During the case retrieval process, the cosine similarity is used to calculate the similarity between cases. The cosine similarity formula is cos(θ)=(A·B) / (||A||·||B||), where A and B are the feature vectors of two cases. For an office area with a building area of 280 square meters and a distance of 450 meters from the subway station, the feature vector obtained after normalization is (0.28, 1, 0, 0, 0.6). Calculate the similarity with all cases in the database, and a total of 52 cases with similarity greater than 0.8 are selected. During the principal component analysis process, the eigenvalues of the feature covariance matrix are calculated to be 2.8, 1.5, 0.9, 0.5, and 0.3 respectively. The corresponding eigenvectors form the principal components. The contribution rate of the first principal component is 47%, the second principal component is 25%, and the third principal component is 18%. The cumulative contribution rate reaches 90%, and these three principal components are retained. The information gain ratio criterion is used to select the splitting feature during the decision tree construction process. Information gain ratio = Information gain / Feature entropy.Taking the floor area as an example, the samples are divided into two groups according to the area: less than 200 square meters and greater than or equal to 200 square meters. The entropy value before splitting is calculated to be 1.8, the entropy value after splitting is 0.9, the information gain is 0.9, the feature entropy is 0.8, and the information gain ratio is obtained as 1.125. The formula for the rule confidence is confidence = the number of samples satisfying the rule / the number of samples in the conditional part of the rule. Taking the rule "If the floor area is greater than 200 square meters and the distance to the subway station is less than 500 meters, then the tax rate is 12%" as an example, the total number of samples satisfying the condition is 40, and among them, 32 samples have a tax rate of 12%. The calculated confidence is 0.8. The formula for the support is support = the number of samples satisfying the rule / the total number of samples, and the support of this rule is 0.15. During the training process of the multi-layer perceptron, the backpropagation algorithm is used to optimize the network parameters, and the cross-entropy function is used as the loss function. For the input samples, the output values of the neurons in each layer are calculated through forward propagation. For the hidden layer, the ReLU activation function: f(x) = max(0, x) is used, and the Softmax function is used in the output layer to convert the output into a probability distribution. After 1000 rounds of training, the prediction accuracy on the validation set reaches 88%. During the correction process of the prediction results, the weight of the usage attribute is set to 0.4, the weight of the area ratio is set to 0.4, and the weight of the usage intensity is set to 0.2. For a mixed-use area, the predicted probability of the office tax rate is 0.65, and the predicted probability of the residential tax rate is 0.35. Considering that the office area ratio of this area is 0.7, the tax rate prediction value is finally adjusted to the office tax rate.
[0029] Step S105, through the logistic regression algorithm, analyze the influence weights of the area and tax rate of each region on the comprehensive tax rate, and use the weighted average method to calculate the comprehensive tax rates of the office area and the residential area, and determine the proportion coefficients of each region in the tax amount calculation.
[0030] Collect the original attribute data of the building area, and process the original attribute data through the min-max normalization method to obtain the first normalized data; calculate the mean and variance of the area data sequence and the tax rate data sequence according to the first normalized data, and process the data sequence through the centering method to obtain the second normalized data; construct a logistic regression predictor based on the second normalized data, and optimize the parameters of the predictor using the stochastic gradient descent method to obtain the area weight value and the tax rate weight value; construct a linear combination equation according to the area weight value and the tax rate weight value, calculate the parameters of the equation using the least squares method to obtain the influence coefficient, normalize the influence coefficient through the sigmoid function to obtain the regional contribution degree, and calculate the proportion coefficient of the tax amount calculation using the additive combination method.
[0031] Specifically, for the office area and residential area, original attribute data such as building area, usable area, actual utilization rate, transportation convenience, decoration grade, equipment and facilities are collected. The original attribute data is normalized using the min-max normalization method to obtain the first normalized data. The mean and variance of the area data sequence and tax rate data sequence are calculated respectively based on the first normalized data, and the data sequence is processed using the centering method to obtain the second normalized data. A logistic regression predictor is constructed based on the second normalized data. The learning rate is set to 0.01 and the number of iterations is set to 1000. The random gradient descent method is used to optimize the predictor parameters to obtain the area weight value and tax rate weight value. A linear combination equation is constructed based on the area weight value and tax rate weight value, and the equation parameters are calculated using the least squares method to obtain the area influence coefficient and tax rate influence coefficient. The comprehensive influence value of each area is calculated based on the area influence coefficient and tax rate influence coefficient, and the sigmoid function is used to normalize the influence value to obtain the area contribution degree. According to the area contribution degree, the area proportion weight and tax rate proportion weight are set, and the additive combination method is used to calculate the comprehensive proportion coefficient of each area, and the final tax calculation proportion coefficient is obtained through normalization. In the tax rate calculation of the office and residential mixed area, the original attribute data shows obvious numerical differences. The building area of the office area is 280 square meters, the usable area is 252 square meters, and the actual utilization rate is 0.9; the building area of the residential area is 160 square meters, the usable area is 144 square meters, and the actual utilization rate is 0.85. The min-max normalization method is used for normalization, and the conversion formula is x'=(x - min) / (max - min). After processing, the area eigenvalue of the office area is 0.7, and the area eigenvalue of the residential area is 0.4. The standardized data is centered, and the mean μ and standard deviation σ of the data sequence are calculated, and the conversion is performed using x'=(x - μ) / σ. Taking the area data as an example, the office area deviates from the mean by 2.1 standard deviations, and the residential area deviates from the mean by 1.2 standard deviations. The converted data is more suitable for logistic regression modeling. The logistic regression predictor uses the random gradient descent method to optimize the parameters, and the loss function is the cross-entropy function, L=-Σ(ylog(p)+(1 - y)log(1 - p)), where y is the actual label and p is the predicted probability. Through 1000 iterations of optimization, the area weight of the office area converges to 0.65, and the tax rate weight converges to 0.35; the area weight of the residential area converges to 0.55, and the tax rate weight converges to 0.45. The linear combination equation is in the form of ax + by + c, where x is the area feature and y is the tax rate feature. The parameter values are calculated using the least squares method, and the equation for the office area is 0.65x + 0.35y + 0.2, and the equation for the residential area is 0.55x + 0.45y + 0.15. Substituting the standardized eigenvalue into the equation, the comprehensive influence value of the office area is calculated to be 0.72, and that of the residential area is 0.58.The sigmoid function σ(x) = 1 / (1 + e^(-x)) is used to normalize the influence values and map them to the interval (0, 1). The normalized influence value for the office area is 0.67, and for the residential area is 0.54. Based on these values, weights are set. The weight for the area proportion is taken as 0.6, and the weight for the tax rate proportion is taken as 0.4. The additive combination w1x1 + w2x2 is used to calculate the comprehensive proportion coefficient. The final calculation result shows that the proportion coefficient of the office area in the tax amount calculation is 0.63, considering its larger area proportion and higher utilization rate; the proportion coefficient of the residential area is 0.37, reflecting its smaller area scale and relatively lower usage intensity. This weighted calculation method based on multi-dimensional features reasonably balances the contribution proportions of different use areas in the tax amount distribution.
[0032] Step S106, integrating the area, tax rate, and proportion coefficient of each region, using the dynamic programming algorithm to establish a tax amount calculation model to obtain the payable tax amount of the building. At the same time, through the sensitivity analysis method, evaluate the influence degree of the changes in the area and tax rate of the region on the tax amount, and form a tax optimization plan through the intelligent tax assistant.
[0033] Construct a tax amount state transition equation f(i, j) including the region number i and the area value j. The state transition equation is calculated using the dynamic programming algorithm to obtain the first optimal solution; according to the first optimal solution, construct an objective function z, which includes the sum of the products of the area adjustment amount xi and the tax rate adjustment amount yi of the region, and solve it using the simplex method to obtain the second optimal solution; calculate the partial derivative of the area change amount with respect to the tax amount for the second optimal solution to obtain the area sensitivity matrix, and calculate the partial derivative of the tax rate change amount with respect to the tax amount to obtain the tax rate sensitivity matrix; construct a constraint equation set according to the area sensitivity matrix and the tax rate sensitivity matrix, solve the optimal adjustment plan using the Lagrange multiplier method, and iteratively optimize the optimal adjustment plan using the gradient descent method to obtain the third optimal solution.
[0034] Specifically, construct a tax amount state transition equation f(i, j) = max{f(i - 1, j), f(i - 1, j - vi) + wi} according to the area data, tax rate data, and proportion coefficient of each region, where i represents the region number, j represents the area value, vi represents the area of the region, and wi represents the tax amount of the region. The first optimal solution is calculated using the dynamic programming algorithm. Set the area change range and tax rate change range as constraint conditions for the first optimal solution, and construct an objective function z = Σ(xi * yi), where xi represents the area adjustment amount of the region and yi represents the tax rate adjustment amount of the region, and solve it using the simplex method to obtain the second optimal solution. Based on the second optimal solution, construct a sensitivity Si calculation formula as , calculate the partial derivative of the change in area with respect to the tax amount to obtain the area sensitivity matrix, and calculate the partial derivative of the change in tax rate with respect to the tax amount to obtain the tax rate sensitivity matrix. Construct a system of constraint equations based on the area sensitivity matrix and the tax rate sensitivity matrix, and use the Lagrange multiplier method to solve the optimal adjustment plan to obtain the area adjustment coefficient and the tax rate adjustment coefficient for each region. Calculate the area adjustment amount and the tax rate adjustment amount for each region based on the adjustment coefficient, set the adjustment step size to 0.01, and use the gradient descent method to iteratively optimize the adjustment amount to obtain the third optimal solution. Construct a tax burden impact evaluation function g(x) = Σ(αixi + βiyi) for the third optimal solution, where αi and βi represent the impact weights of area and tax rate respectively, and calculate the impact value of each adjustment amount on the overall tax burden by taking the derivative of the function. When solving the tax amount optimization problem by dynamic programming, take the office area and residential area of a building as an example. The office area is 280 square meters, the tax rate is 12%, and the proportion coefficient is 0.65; the residential area is 160 square meters, the tax rate is 4%, and the proportion coefficient is 0.35. Construct a state transition equation f(i,j), where i represents the current region number, j represents the current area value, vi represents the area of region i, and wi represents the tax contribution of region i. Obtain the state matrix through state recursion. In the matrix, f(1,280) = 21.84 represents the optimal tax amount when only considering the office area, and f(2,440) = 27.44 represents the optimal tax amount when considering all regions. Set the area change range to ±20 square meters and the tax rate change range to ±2 percentage points, and construct a linear programming model z = Σ(xi*yi), where z represents the total tax amount, xi represents the area adjustment amount, and yi represents the tax rate adjustment amount. Take the partial derivative of the objective function to calculate the sensitivity. Taking the office area as an example, the area sensitivity means that for every 1 square meter increase in area, the tax amount increases by 0.078 million yuan, and the tax rate sensitivity It means that for every 1 percentage point increase in the tax rate, the tax amount increases by 23,400 yuan. The area sensitivity of the residential area is 0.014, and the tax rate sensitivity is 0.64, indicating that the impact of tax rate adjustment on the tax amount is more significant. Based on the sensitivity analysis results, the Lagrangian function L(x,y,λ)=z(x,y)+λ(g(x,y)-c) is constructed, where g(x,y) represents the constraint condition, c represents the constraint boundary, and z(x,y) is the original objective function. Solving the Lagrangian equation gives the optimal adjustment plan: the area of the office area is reduced by 15 square meters, and the tax rate is reduced by 1.2 percentage points; the area of the residential area is increased by 12 square meters, and the tax rate is increased by 0.8 percentage points. The gradient descent method is used to optimize the adjustment plan, with the learning rate set to 0.01, and the iteration stops when the gradient value is less than 0.001. After 78 rounds of iteration, the final plan is obtained: the area of the office area is adjusted to 267 square meters, and the tax rate is adjusted to 10.9%; the area of the residential area is adjusted to 170 square meters, and the tax rate is adjusted to 4.7%. The impact evaluation function g(x)=Σ(αixi+βiyi) is constructed, where αi and βi are 0.6 and 0.4 respectively. The calculated impact value of the area adjustment on the tax burden is -0.86, the impact value of the tax rate adjustment is -1.24, and the comprehensive impact value is -2.1, indicating that the optimization plan realizes the optimization of the tax burden structure on the premise of keeping the total tax amount basically stable.
[0035] Step S107: Generate an intelligent tax advice report in the intelligent tax assistant according to the comprehensive tax rate, payable tax amount and tax optimization plan of the building.
[0036] The association rule mining algorithm is used to calculate the comprehensive tax rate index of the building, and the first association rule set is obtained according to the support degree and confidence degree between the indexes; hierarchical clustering processing is performed on the first association rule set, and the rules are grouped by the clustering quantity parameter to obtain the first report framework, which includes background information, data analysis, optimization suggestions and conclusion summary; keyword groups are extracted according to the first report framework, the semantic similarity is calculated by the word vector embedding method, and the semantic clustering is performed by the cosine distance method to obtain the second report framework; cases are retrieved from the historical report library for the second report framework, and cases within the similarity threshold are selected by the text similarity calculation method to obtain a standardized report template, which includes a professional term dictionary and a numerical index rule library.
[0037] Specifically, data indicators for the comprehensive tax rate, payable tax amount, and tax optimization plan of a building are collected. The support and confidence between the indicators are calculated using the association rule mining algorithm. The minimum support threshold is set to 0.3, and the minimum confidence threshold is set to 0.7, resulting in the first association rule set. Based on the first association rule set, a report title system is constructed. The hierarchical clustering method is used to group the rules, and the number of clusters is set to 4, corresponding to the background information, data analysis, optimization suggestions, and conclusion summary in the report, obtaining the first report framework. According to the first report framework, keyword groups and syntactic structures are extracted. The word vector embedding method is used to calculate semantic similarity, and the cosine distance method is used to cluster the semantics, obtaining the second report framework. Based on the second report framework, similar cases are retrieved from the historical report library. The text similarity calculation method is used to select cases with a similarity greater than 0.8 as reference templates, obtaining the first report template. For the first report template, a professional term dictionary and a numerical indicator rule library are constructed. The string matching method is used to identify professional terms in the text, and the regular expression method is used to extract numerical indicators, obtaining the second report template. Based on the second report template, standardized text content is generated. The text verification rules are used to check the numerical accuracy, unit annotation, and format specification, and the structured report content is formed through the natural language generation method. During the generation of the intelligent tax advice report, the association rule mining algorithm is used to analyze the data indicators, and multiple groups of association rules are found. When the office area is greater than 200 square meters, the support for the comprehensive tax rate exceeding 8% is 0.35, and the confidence is 0.82; when the residential area is greater than 150 square meters, the support for the tax rate optimization space being greater than 1 percentage point is 0.42, and the confidence is 0.75. The report framework is constructed through the hierarchical clustering method, and the content with similar themes is grouped under the same title. The data analysis part includes three sub-themes: area distribution, tax rate structure, and space utilization rate. The optimization suggestions part includes three sub-themes: area adjustment suggestions, tax rate optimization measures, and comprehensive benefit evaluation. The background information and conclusion summary are used as independent themes respectively. The semantics of the report are vectorized, and the word vector embedding method is used to calculate semantic similarity. Taking "tax rate optimization" as an example, the similarity with "tax burden adjustment" is 0.92, the similarity with "tax planning" is 0.88, and the similarity with "tax amount calculation" is 0.75. Based on the semantic similarity clustering results, similar concepts are uniformly expressed to form a standardized professional term system. When retrieving similar cases from the historical report library, the cosine distance method is used to calculate the text similarity. For a new report text, its similarity distribution with historical cases is 0.85 for case 1, 0.82 for case 2, and 0.76 for case 3. Cases 1 and 2 with a similarity greater than 0.8 are selected as reference templates, and the expression methods and format specifications therein are extracted.Define standard terminology expressions in the professional term dictionary, such as "building area" is uniformly expressed as "taxable area", "tax rate" is uniformly expressed as "applicable tax rate", and "optimization plan" is uniformly expressed as "tax burden optimization suggestion". The numerical indicator rules stipulate that the area is retained as an integer, the tax rate is retained to 1 decimal place, the amount is retained to 2 decimal places, and the proportion is retained to 1 place before the percentage sign. The text standardization check includes multiple dimensions. The numerical accuracy check found that the area data "280.45 square meters" should be corrected to "280 square meters"; the unit marking check found that the percentage sign was missing after "12.5%"; the format standardization check found that the title level was inconsistent, and the third-level title "3.2.1 Tax rate structure analysis" should be adjusted to "3.2 Tax rate structure analysis". Through the natural language generation method, the standardized content is organized into a structured report to form a tax recommendation report with rigorous logic and standardized expression.
[0038] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential features of the present application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present application. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. An intelligent tax assistant interaction method, characterized in that: The method comprises: Through laser scanning and panoramic cameras, the spatial layout diagram and area data of each area after the internal partition transformation of the building are obtained, image segmentation and target detection are performed on the spatial layout diagram, and the geometric features and spatial relationships of each area are extracted. In combination with the building design specifications and tax classification standards, the use nature of each area is determined, and the use nature includes office, residential and mixed use. Specifically, a three-dimensional laser scanner is used to obtain the first point cloud data inside the building, and the first texture point cloud model is obtained by registering the first point cloud data with the panoramic image data; Performing spatial clustering processing on the first texture point cloud model to obtain a second texture point cloud model, extracting wall surfaces using a region growing algorithm based on the second texture point cloud model and calculating the relative position relationship between walls to obtain a first spatial connectivity graph; Performing vectorization processing on the first spatial connectivity graph to obtain a first plane layout graph, and performing image segmentation processing on the first plane layout graph to obtain a second plane layout graph; Extracting shape feature vectors and space topology feature vectors from the second plan layout diagram, using a support vector machine algorithm to train the shape feature vectors and space topology feature vectors to obtain a room use classification model, and labeling the rooms for office and residential use according to the classification model; If the intelligent tax assistant determines that the area is used for office or residential purposes, it uses a supervised machine learning algorithm to train the existing tax classification sample data and establish a tax classification model to determine the tax rate category corresponding to the office area and the residential area; If the intelligent tax assistant determines that the use nature of the area is mixed, the mixed nature area is subdivided into several sub-areas through the cluster analysis algorithm, and the sub-areas are determined to be office or residential based on the area proportion and use characteristics of the sub-areas. Specifically, the first feature vector is obtained by collecting the basic feature data of the door and window positions, partition walls and channel connections in the mixed nature area and using the maximum and minimum value normalization method; Calculating the Euclidean distance according to the first eigenvector to obtain a first similarity matrix, and using a density clustering algorithm to obtain a first sub-region division result; Extracting the room connectivity relationship based on the first sub-region division result, obtaining a second similarity matrix by calculating the channel width and connectivity between adjacent rooms, and obtaining the second sub-region division result using a graph cut algorithm; Extracting building area data according to the second sub-area division result, collecting crowd density data through infrared sensors, using a hierarchical clustering algorithm to obtain office clusters and residential clusters, and determining the use nature of the sub-area according to the area weight and cluster feature weight; Through case-based reasoning algorithms, similar cases to the sub-region are retrieved from the existing tax case database, and relevant attributes, including area, usage function, and occupancy, are compared and analyzed to predict the tax rate tendency of the region and form preliminary tax recommendations; The logistic regression algorithm is used to analyze the influence of each area and tax rate on the comprehensive tax rate, and the weighted average method is used to calculate the comprehensive tax rate of office and residential areas to determine the proportion coefficient of each area in the tax calculation; Based on the area, tax rate and proportion coefficient of each area, a dynamic programming algorithm is used to establish a tax calculation model to obtain the tax payable for the building. At the same time, the sensitivity analysis method is used to evaluate the impact of changes in regional area and tax rate on the tax amount, and a tax optimization plan is formed through the intelligent tax assistant; Generate an intelligent tax recommendation report in the Intelligent Tax Assistant based on the building’s comprehensive tax rate, tax payable, and tax optimization plan.
2. The method according to claim 1, characterized in that If the intelligent tax assistant determines that the use nature of the area is office or residential, a supervised machine learning algorithm is used to train the existing tax classification sample data to establish a tax classification model for determining the tax rate category corresponding to the office area and the residential area, including: The standard deviation method is used to identify outliers in the tax classification sample data, and the mean filling method is used to supplement the missing values according to the outlier identification results to obtain the first sample data; Perform minimum and maximum value normalization processing on the first sample data, and perform one-hot encoding conversion on the normalization processing result to obtain third sample data; A variance selection method is used to remove low-discrimination features from the third sample data, and a principal component analysis method is used to extract the principal component whose cumulative contribution rate reaches a specified threshold value from the data after the low-discrimination features are removed to obtain a sixth sample data; A random forest classifier is established based on the sixth sample data, and the probability value of the tax rate category is calculated by the random forest classifier. If the probability value is greater than a specified threshold, the corresponding tax rate category is adopted; if the probability value is less than the specified threshold, the tax rate category is determined by a probability weighted method.
3. The method according to claim 1, characterized in that: The case-based reasoning algorithm retrieves cases similar to the sub-region from the existing tax case database, compares and analyzes relevant attributes, including area, usage function, and occupancy, predicts the tax rate tendency of the area, and forms preliminary tax recommendations, including: The first attribute data of building area, use function and traffic convenience are extracted from the tax case database, and the feature vector is obtained through the minimum and maximum value standardization method and the one-hot encoding method; Retrieving cases with similarity greater than a threshold from a tax case database using a cosine similarity method according to the feature vector to obtain a first case set; Extracting the second attribute data of regional area ratio, space utilization rate and personnel density from the first case set, and selecting the principal component with the cumulative contribution rate meeting the requirements by the principal component analysis method to obtain the second case set; A decision tree model was constructed based on the second case set, the information gain ratio was used to select the optimal split feature extraction decision path, and the tax rate prediction model was obtained through multi-layer perceptron training.
4. The method according to claim 1, characterized in that: The logistic regression algorithm is used to analyze the weight of the impact of each area and tax rate on the comprehensive tax rate, and the weighted average method is used to calculate the comprehensive tax rate of the office area and the residential area, and the proportion coefficient of each area in the tax calculation is determined, including: Collecting original attribute data of the building area, and processing the original attribute data by a minimum-maximum value normalization method to obtain first standardized data; Calculating the mean variance of the area data sequence and the tax rate data sequence according to the first standardized data, and processing the data sequence using a centralization method to obtain second standardized data; Constructing a logistic regression predictor based on the second standardized data, and optimizing the predictor parameters by using a stochastic gradient descent method to obtain an area weight value and a tax rate weight value; A linear combination equation is constructed based on the area weight value and the tax rate weight value, the least squares method is used to calculate the parameters of the equation to obtain the influence coefficient, the influence coefficient is normalized by the sigmoid function to obtain the regional contribution, and the additive combination method is used to calculate the tax amount calculation ratio coefficient.
5. The method according to claim 1, characterized in that The tax calculation model is established by integrating the area, tax rate and proportion coefficient of each area, using a dynamic programming algorithm to obtain the tax payable for the building. At the same time, the sensitivity analysis method is used to evaluate the impact of changes in regional area and tax rate on the tax amount, and a tax optimization plan is formed through the intelligent tax assistant, including: Constructing a tax amount state transfer equation f(i, j) including the area number i and the area value j, wherein the state transfer equation is calculated using a dynamic programming algorithm to obtain a first optimal solution; Constructing an objective function z according to the first optimal solution, the objective function includes the product sum of the regional area adjustment amount xi and the tax rate adjustment amount yi, and using the simplex method to solve to obtain the second optimal solution; For the second optimal solution, the partial derivative of the area change with respect to the tax amount is calculated to obtain an area sensitivity matrix, and the partial derivative of the tax rate change with respect to the tax amount is calculated to obtain a tax rate sensitivity matrix; A set of constraint equations is constructed according to the area sensitivity matrix and the tax rate sensitivity matrix, an optimal adjustment scheme is solved by using the Lagrange multiplier method, and the optimal adjustment scheme is iteratively optimized by using the gradient descent method to obtain a third optimal solution.
6. The method according to claim 1, characterized in that According to the building's comprehensive tax rate, tax payable and tax optimization plan, an intelligent tax recommendation report is generated in the intelligent tax assistant, including: The association rule mining algorithm is used to calculate the comprehensive tax rate index of buildings, and the first association rule set is obtained according to the support and confidence between the indicators; Performing hierarchical clustering processing on the first association rule set, grouping the rules by cluster quantity parameters to obtain a first report framework, wherein the first report framework includes background information, data analysis and optimization suggestions, and conclusion summary; Extract key words according to the first report framework, calculate semantic similarity by word vector embedding method, and use cosine distance method to cluster semantics to obtain the second report framework; For the second report framework, cases are retrieved from the historical report library, and cases within a similarity threshold are selected through a text similarity calculation method to obtain a standardized report template, wherein the report template includes a professional terminology dictionary and a numerical indicator rule library.
Citation Information
Patent Citations
Point cloud plane identification and edge detection method
CN116012399A
Method for quickly constructing three-dimensional block model of building
CN117523125A