House type drawing intelligent identification method and device based on image recognition

By combining the OpenPose library, ResNet module, and OCR module, along with fixed-frequency acoustic echo measurement technology, the problems of complex structures and low-quality image processing in floor plan recognition are solved. This enables accurate measurement of floor plan space and text recognition, improving the overall accuracy and efficiency of intelligent floor plan recognition.

CN120726641BActive Publication Date: 2025-11-18BBMG TIANTAN FURNITURE CO LTD +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511189834.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-11-18
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing intelligent floor plan recognition technology is not accurate enough when recognizing complex structures, has difficulty processing low-quality images, lacks the ability to process video information, cannot accurately measure space dimensions, and has low text recognition accuracy, which affects interior design and space planning.

Method used

The OpenPose library is used for key point recognition, combined with the ResNet module for heat map analysis, the OCR module is used to extract text information, and the scale is marked by the scale module. The spatial dimensions are measured by combining the fixed frequency sound wave echo measurement technology, and a manual verification mechanism is introduced to ensure accuracy.

Benefits of technology

It significantly improves the accuracy of key structure recognition in floor plans, enhances the ability to process low-quality images, enables accurate measurement of floor plan spaces and text recognition, expands application scenarios, and improves work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726641B_ABST
    Figure CN120726641B_ABST
Patent Text Reader

Abstract

The application discloses a house type drawing intelligent identification method and device based on image recognition, and relates to the technical field of house type drawing identification.The application comprises the following steps: a collection module obtains original data from a mobile terminal, the original data comprising image information or video information, the collection module forwards the original data to a database for storage, an OpenPose library obtains the original data or feedback information from the database, executes a key point identification program to obtain a house type framework, the OpenPose library transmits the house type framework to a ResNet module, the ResNet module obtains the original data or feedback information from the database, executes a hotspot map analysis program to obtain identification data, the ResNet module combines the identification data and the house type framework into first data and transmits the first data to an OCR module.The application realizes intelligent identification of house type drawings through multi-module collaborative processing, improves identification accuracy and efficiency, and solves the technical problems of low accuracy and poor efficiency of traditional house type drawing identification methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of floor plan recognition technology, specifically to a method and apparatus for intelligent floor plan recognition based on image recognition. Background Technology

[0002] In the field of floor plan generation, existing technologies are mainly divided into two categories: traditional manual drawing and semi-automatic recognition based on image recognition. Traditional methods rely on manual on-site measurement (such as using measuring tapes or laser rangefinders) and manual drawing using CAD software. Data such as wall lengths and door / window positions are manually recorded and converted into graphical representations. With the development of computer vision technology, image recognition-based assisted methods have emerged. These methods use edge detection, feature extraction, and other algorithms to identify elements such as walls, doors, and windows from building images. Some methods integrate OCR text recognition technology to extract room names, assisting in the initial drawing of the floor plan.

[0003] Patent CN110189398A discloses a method for generating floor plans based on indoor images. This method identifies an initial image using a preset image recognition model, determines whether it is an indoor image, detects floor plan generation components within the indoor image, and generates a 3D floor plan based on these components and their position and dimensions. However, existing intelligent floor plan recognition technologies still have some problems and shortcomings: Existing technologies lack accuracy in recognizing key structures such as walls, doors, and windows in floor plans, especially for complex floor plans, where the recognition accuracy is low and difficult to meet practical application needs. Furthermore, for floor plans with poor image quality... For images that are blurry, distorted, or poorly lit, the recognition accuracy of existing technologies is significantly reduced, making it impossible to effectively process floor plans under various quality conditions. Existing technologies mainly process static images and lack the ability to effectively extract and process floor plan data from video information, limiting the application scenarios of the technology. Existing floor plan recognition methods lack accurate measurement methods for the actual dimensions of the floor plan space, making it difficult to provide accurate spatial dimension information, which affects subsequent decoration design and space planning. Existing technologies are insufficient in extracting text information, especially for handwritten annotations or special fonts, where the recognition accuracy is low, affecting the integrity and usability of the floor plan.

[0004] Therefore, there is an urgent need for an intelligent floor plan recognition method that can accurately identify key structures in floor plans, adapt to different image qualities, process video information, accurately measure spatial dimensions, and precisely extract text information to solve existing technical problems. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a method and apparatus for intelligent recognition of floor plans based on image recognition, which solves the problems mentioned in the background section.

[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent floor plan recognition method based on image recognition, comprising the following steps:

[0007] Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage.

[0008] Step 2: The OpenPose library retrieves raw data or feedback information from the database, executes the key point recognition program to obtain the floor plan framework, and then transfers the floor plan framework to the ResNet module.

[0009] Step 3: The ResNet module retrieves raw data or feedback information from the database, executes the heat map analysis program to obtain the identification data, and merges the identification data and the floor plan framework into the first data and transmits the first data to the OCR module.

[0010] Step 4: The OCR module retrieves raw data or feedback information from the database, executes the text extraction program to obtain scanned data, and merges the first data and scanned data into second data and transmits the second data to the ruler module.

[0011] Step 5: The scale module marks the scale in the second data to obtain the floor plan, and the scale module transmits the floor plan to the verification module;

[0012] Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which then outputs the floor plan to the mobile terminal or display platform for users to view. If the verification fails, the verification module marks the feedback information in the floor plan, uploads the feedback information to the database, and re-executes Step 2.

[0013] Furthermore, the key point recognition procedure specifically includes the following steps:

[0014] Step 201: The OpenPose library determines whether the raw data is image information or video information. If it is image information, it executes the image analysis process; if it is video information, it executes the video analysis process.

[0015] The video analytics process executes in the following steps:

[0016] Step 208: Separate the picture information and audio information from the video information. The picture information consists of consecutive frames in time sequence. Each consecutive frame is consistent with the image information. Since the picture information has a larger original data than the image information, it has enough feature data for analysis. Therefore, extract several typical frames from the picture information and directly perform key point analysis. Calculate the picture change rate between adjacent consecutive frames pixel by pixel. Calculate the percentage of pixels whose pixel values ​​change out of the total number of pixels. Set consecutive frames with a picture change rate greater than or equal to 30% as typical frames.

[0017] Step 209: Establish the second key point model. The typical frame retains the RGB three channels as continuous features and inputs them into the second key point model. The first key point model is based on the VGG19 convolutional neural network. The second key point model performs row convolution training on the continuous features. If there is feedback information in the database and the output of the convolution training is the same as the feedback information, the convolution training is performed again. After the training is completed, the continuous key points and continuous edge features in the typical frame are marked.

[0018] In this process, when the mobile terminal is shooting video information, it uses its own speaker to play a fixed-frequency sound wave. The fixed-frequency sound wave ranges from 500Hz to 2kHz. The fixed-frequency sound wave spreads outward, hits the wall, and bounces back to the mobile terminal. The microphone of the mobile terminal then obtains the audio information. Essentially, the mobile terminal uses the spread echo of the fixed-frequency sound wave to analyze the interior space of the apartment.

[0019] Step 210: Execute the echo measurement process to obtain the distance d from the mobile terminal to the wall in a single direction;

[0020] Step 211: Repeat the echo measurement process to calculate the distance d in multiple directions. Since the microphone used by the mobile terminal is a single microphone, the collected audio information is not directional and the specific direction of the distance d cannot be located. The distances d in multiple directions are arranged and combined to obtain several audio spaces. At this time, the spatial dimension cannot determine the specific frame and output multiple possible results. Then, the continuous key points and continuous edge features output by the subsequent second key point model are verified.

[0021] Step 212: Cross-validate the continuous keypoints and continuous edge features from Step 209 with several audio spaces from Step 211. Mark the audio space with the highest overlap rate with the continuous keypoints and continuous edge features as the audio feature, delete the remaining audio spaces, and use the audio feature as the basis. The audio feature is calculated based on fixed-frequency sound waves, and the calculation results regarding the wall distance are more accurate and reliable. Correct the continuous keypoints and continuous edge features to obtain the house frame. That is, when there is an error between the distance of the continuous keypoints and continuous edge features and the audio feature, the distance of the audio feature shall prevail.

[0022] Furthermore, the heatmap analysis program specifically includes the following steps:

[0023] Step 301: The ResNet module adjusts the pixel size of the raw data and normalizes the pixel values ​​of the raw data. If the raw data is video information rather than image information, the first frame of the video information is used as the input raw data.

[0024] Step 302: Load the pre-trained ResNet model, such as ResNet-50, and fine-tune it based on the apartment type identification dataset. The apartment type identification dataset includes labeled data such as single doors, double doors, sliding doors, and windows. The data is manually labeled in advance by data labelers. The apartment type identification dataset is stored in the database. The output layer uses the Sigmoid activation function to make the pixel value range of each channel [0,1], thus obtaining a multi-channel heat map. If there is feedback information in the database and the heat map output result is the same as the feedback information, the heat map is re-output.

[0025] Step 303: Set a confidence threshold T for the heatmap of each channel, for example, T=0.6, retain regions with pixel values ​​greater than the confidence threshold T, filter out low-probability noise, and obtain a binarized mask. The filtering formula is as follows: Mc is the mask of the identifier, which determines whether the identifier is retained at the pixel coordinates (x, y). (x, y) is the pixel coordinate. Hc(x, y) is the heat map of the identifier. The original heat map contains noise and redundant candidate regions. Post-processing can filter out the identifier regions with high confidence and clear boundaries, thereby improving the accuracy of subsequent recognition.

[0026] Step 304: Match the heat map of each channel after filtering in step 303 with the house frame for location verification. The house frame is the wall orientation output by the OpenPose library. Delete the heat map that fails the location verification to ensure that the identifier is consistent with the spatial logic of the house frame, so as to provide reliable information for subsequent merging of the first data.

[0027] Step 305: Summarize the heat maps verified in Step 304 into identification data.

[0028] Furthermore, the text extraction procedure specifically includes the following steps:

[0029] Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the original data, improves the contrast between text and background through adaptive histogram equalization, and then binarizes the original data to obtain a binary image, which further separates the text and background. The Otsu algorithm is used to automatically determine the binarization threshold. If the original data is video information rather than image information, the first frame of the video information is used as the input original data.

[0030] Step 402: Set the scanning window. The length and width of the scanning window are both 10% of the original data. Based on TesseractOCRAPI, use the scanning window to scan the binary image line by line and extract the text from the binary image to obtain the text data. If there is feedback information in the database and the output text data is the same as the feedback information, repeat step 402 to output the text data.

[0031] Step 403: Establish a semantic similarity model. The semantic similarity model is based on paraphrase-multilingual-MiniLM-L12-v2 in the sentence-transformers library. It is used to calculate the similarity between text vectors and to match misspelled words with similar shapes or meanings. Input the text data into the semantic similarity model for error correction. After the error correction is completed, the scanned data is obtained.

[0032] Furthermore, marking the scale specifically includes the following steps:

[0033] Step 501: The ruler module counts the pixel size data of the label data, such as the number of pixels occupied by the length and width of different doors and windows in the label data. The average of all pixel size data is used to obtain the pixel ratio, which is the ratio of the actual size of a single pixel.

[0034] Step 502: Set the scale bar. Multiply the actual size corresponding to the pixel ratio by 20 and assign the value to the scale bar, then mark it in the second data.

[0035] Furthermore, the verification module verifies the floor plan specifically by including the following steps:

[0036] Step 601: The verification module is a manual data backend. After receiving the floor plan, the verification module synchronously retrieves the corresponding raw data from the database.

[0037] Step 602: The manual data backend identifies the differences between the original data and the floor plan to determine whether the floor plan matches the original data;

[0038] Step 603: If the original data matches the floor plan, it is marked as verified. If the original data does not match the floor plan, the manual data editor will mark the location or area on the floor plan that does not match the original data as feedback information.

[0039] Furthermore, the image analysis process executes the following steps:

[0040] Step 202: Convert the image information into a grayscale image, and then perform Gaussian blur to remove noise from the grayscale image. The kernel size is adjusted according to the noise level of the image information.

[0041] Step 203: Perform Canny operator edge detection on the denoised grayscale image. Set dual thresholds to control edge sensitivity and distinguish grayscale images. The low threshold is used to retain weak edges, and the high threshold is used to identify strong edges. The ratio of the low threshold area to the high threshold area is controlled to be 1:3. The specific values ​​of the dual thresholds are dynamically adjusted according to the ratio of the low threshold area to the high threshold area.

[0042] Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image. Merge the weak edges of the Canny operator with the horizontal or vertical edges of the Sobel operator, take the union, and use it to enhance the weak edges. Use the image processing method of dilation followed by erosion to connect the broken edges to obtain the total edge map. The broken edges include weak edges, strong edges, horizontal edges, and vertical edges.

[0043] Step 205: Use Hough transform to extract line segments from the total edge map. The extracted line segments are potential wall lines or door frame lines.

[0044] It should be noted that, since the feature data carried by the image information is limited, steps 202 to 205 are required to preprocess the image information and extract the edge features in the image information. The edge features refer to the straight line segments in the total edge map, thereby improving the accuracy of subsequent key point recognition.

[0045] Step 206: Establish the first key point model. The image information retains the RGB three channels as the basic features and is input into the first key point model. The extracted line segments are then input into the first key point model as auxiliary features. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features. After training, the key points in the image information are marked. The key points are the intersection points of the edge line segments in the floor plan.

[0046] Step 207: Cross-validate the line segments from Step 205 with the key points from Step 206 by overlapping them and setting an overlap threshold. If the overlap rate between the intersection points of the line segments and the key points is greater than or equal to the overlap threshold, it means that the recognition rate of the line segments and key points can meet the requirements for forming the house frame. Combine the line segments and key points to obtain the house frame, and the image analysis process ends. If the overlap rate between the intersection points of the line segments and the key points is less than the overlap threshold, it means that the recognition rate of the line segments and key points cannot meet the requirements for forming the house frame. Jump to Step 203 and repeat Steps 203 to 206 until the overlap rate between the intersection points of the line segments and the key points is greater than or equal to the overlap threshold.

[0047] Furthermore, the echo measurement process specifically includes the following steps:

[0048] To construct a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, and then apply the formula... The slope k of the frequency response curve is calculated in real time. f1 is the frequency at any point on the frequency response curve, A1 is the amplitude at any point on the frequency response curve, f2 is the frequency of the next adjacent point of f1 on the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 on the frequency response curve. When the slope k is greater than or equal to 0.5, it means that the frequency response curve increases sharply at that moment, that is, the mobile terminal receives the echo of the fixed-frequency sound wave hitting the wall. The moment corresponding to the slope k being greater than or equal to 0.5 is marked as the echo point t2. The audio information stores the time t1 when the mobile terminal emits the fixed-frequency sound wave. Subtracting the time t1 from the echo point t2 gives the time difference Δt of the fixed-frequency sound wave's round-trip propagation. According to the formula... The distance d from the moving terminal to the wall in a single direction is obtained, where v is the speed of sound under standard atmospheric pressure.

[0049] The intelligent floor plan recognition device based on image recognition includes a data acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module, and an output module. The output of the data acquisition module is connected to the input of the database. The output of the database is connected to the inputs of the OpenPose library, the ResNet module, and the OCR module. The output of the OpenPose library is connected to the input of the ResNet module. The output of the ResNet module is connected to the input of the OCR module. The output of the OCR module is connected to the input of the ruler module. The output of the ruler module is connected to the input of the verification module. The output of the verification module is connected to the input of the output module. The port of the verification module establishes bidirectional communication with the port of the database.

[0050] The acquisition module is used to acquire raw data from the mobile terminal; the OpenPose library is used to identify the floor plan framework of the raw data in the database; the ResNet module is used to identify the identifier data of the raw data; the OCR module is used to extract the scanned data from the raw data; the scale module is used to set the scale; the verification module is used to verify whether the floor plan matches the raw data; the database is used to store the raw data, the feedback information from the verification module, and the computer execution program instructions; and the output module is used to output the floor plan that has passed verification.

[0051] Furthermore, the mobile terminal includes a speaker, an image sensor, and a microphone. The speaker is used to generate fixed-frequency sound waves, the microphone is used to acquire the sound waves that bounce back after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or video information.

[0052] The beneficial effects of this invention are as follows: By employing the OpenPose library for key point recognition and distinguishing between image and video processing workflows, and combining various image processing techniques such as grayscale conversion, Gaussian blur denoising, and Canny edge detection, the accuracy of floor plan recognition is significantly improved. It can accurately identify key structures such as walls, doors, and windows, while also enhancing the processing capability for low-quality images. It innovatively applies fixed-frequency acoustic echo measurement technology to video information processing, achieving accurate measurement of the actual dimensions of the floor plan space and expanding application scenarios. Through ResNet model-based heatmap analysis, the recognition accuracy of floor plan elements such as doors and windows is improved. The OCR module, combined with a semantic similarity model, improves the accuracy of text recognition. The scale module establishes a scale by calculating pixel ratios, achieving accurate measurement of the actual dimensions of the floor plan space. A manual verification mechanism ensures the quality and reliability of the floor plan. Overall, it achieves automated conversion from raw data to standard floor plans, improving work efficiency and demonstrating significant advantages compared to existing technologies.

[0053] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0054] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a system block diagram of the intelligent floor plan recognition device based on image recognition according to the present invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] Please see Figure 1 This invention provides a technical solution: an intelligent floor plan recognition method based on image recognition, comprising the following steps:

[0058] Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage.

[0059] Step 2: The OpenPose library retrieves raw data or feedback information from the database, executes the key point recognition program to obtain the floor plan framework, and then transfers the floor plan framework to the ResNet module.

[0060] Step 3: The ResNet module retrieves raw data or feedback information from the database, executes the heat map analysis program to obtain the identification data, and merges the identification data and the floor plan framework into the first data and transmits the first data to the OCR module.

[0061] Step 4: The OCR module retrieves raw data or feedback information from the database, executes the text extraction program to obtain scanned data, and merges the first data and scanned data into second data and transmits the second data to the ruler module.

[0062] Step 5: The scale module marks the scale in the second data to obtain the floor plan, and the scale module transmits the floor plan to the verification module;

[0063] Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which then outputs the floor plan to the mobile terminal or display platform for users to view. If the verification fails, the verification module marks the feedback information in the floor plan, uploads the feedback information to the database, and re-executes Step 2.

[0064] The key point recognition process specifically includes the following steps:

[0065] Step 201: The OpenPose library determines whether the raw data is image information or video information. If it is image information, it executes the image analysis process; if it is video information, it executes the video analysis process.

[0066] The video analytics process executes in the following steps:

[0067] Step 208: Separate the picture information and audio information from the video information. The picture information consists of consecutive frames in time sequence. Each consecutive frame is consistent with the image information. Since the picture information has a larger original data than the image information, it has enough feature data for analysis. Therefore, extract several typical frames from the picture information and directly perform key point analysis. Calculate the picture change rate between adjacent consecutive frames pixel by pixel. Calculate the percentage of pixels whose pixel values ​​change out of the total number of pixels. Set consecutive frames with a picture change rate greater than or equal to 30% as typical frames.

[0068] Step 209: Establish the second key point model. The typical frame retains the RGB three channels as continuous features and inputs them into the second key point model. The first key point model is based on the VGG19 convolutional neural network. The second key point model is trained by row convolution on the continuous features. After training, the continuous key points and continuous edge features in the typical frame are marked.

[0069] In this process, when the mobile terminal is shooting video information, it uses its own speaker to play a fixed-frequency sound wave. The fixed-frequency sound wave ranges from 500Hz to 2kHz. The fixed-frequency sound wave spreads outward, hits the wall, and bounces back to the mobile terminal. The microphone of the mobile terminal then obtains the audio information. Essentially, the mobile terminal uses the spread echo of the fixed-frequency sound wave to analyze the interior space of the apartment.

[0070] Step 210: Execute the echo measurement process to obtain the distance d from the mobile terminal to the wall in a single direction;

[0071] Step 211: Repeat the echo measurement process to calculate the distance d in multiple directions. Since the microphone used by the mobile terminal is a single microphone, the collected audio information is not directional and the specific direction of the distance d cannot be located. The distances d in multiple directions are arranged and combined to obtain several audio spaces. At this time, the spatial dimension cannot determine the specific frame and output multiple possible results. Then, the continuous key points and continuous edge features output by the subsequent second key point model are verified.

[0072] Step 212: Cross-validate the continuous keypoints and continuous edge features from Step 209 with several audio spaces from Step 211. Mark the audio space with the highest overlap rate with the continuous keypoints and continuous edge features as the audio feature, delete the remaining audio spaces, and use the audio feature as the basis. The audio feature is calculated based on fixed-frequency sound waves, and the calculation results regarding the wall distance are more accurate and reliable. Correct the continuous keypoints and continuous edge features to obtain the house frame. That is, when there is an error between the distance of the continuous keypoints and continuous edge features and the audio feature, the distance of the audio feature shall prevail.

[0073] The heat map analysis procedure specifically includes the following steps:

[0074] Step 301: The ResNet module adjusts the pixel size of the raw data, uniformly modifying it to 224×224 pixels to adapt to the default input of ResNet-50. It normalizes the pixel values ​​of the raw data, converting the pixel values ​​from the range of [0,255] to [0,1] proportionally. If the raw data is video information rather than image information, the first frame of the video information is used as the raw data of the input.

[0075] Step 302: Load the pre-trained ResNet model, such as ResNet-50, and fine-tune it based on the apartment type identification dataset. The apartment type identification dataset includes labeled data such as single doors, double doors, sliding doors, and windows. The data is manually labeled in advance by data labelers. The apartment type identification dataset is stored in the database. The output layer uses the Sigmoid activation function to make the pixel value range of each channel [0,1], thus obtaining a multi-channel heat map.

[0076] Step 303: Set a confidence threshold T for the heatmap of each channel, for example, T=0.6, retain regions with pixel values ​​greater than the confidence threshold T, filter out low-probability noise, and obtain a binarized mask. The filtering formula is as follows: Mc is the mask of the identifier, which determines whether the identifier is retained at the pixel coordinates (x, y). (x, y) is the pixel coordinate. Hc(x, y) is the heat map of the identifier. The original heat map contains noise and redundant candidate regions. Post-processing can filter out the identifier regions with high confidence and clear boundaries, thereby improving the accuracy of subsequent recognition.

[0077] Step 304: Match the heat map of each channel after filtering in step 303 with the house frame for location verification. The house frame is the wall orientation output by the OpenPose library. For example, doors and windows must be located on the walls. Delete the heat map that fails the location verification to ensure that the identifier is consistent with the spatial logic of the house frame, so as to provide reliable information for subsequent merging of the first data.

[0078] Step 305: Summarize the heat maps verified in Step 304 into identification data.

[0079] The text extraction process specifically includes the following steps:

[0080] Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the original data. The command is blurred=GaussianBlur(gray,(5,5),0). Adaptive histogram equalization is used to improve the contrast between text and background. The command is clahe=CLAHE(clipLimit=2.0,tileGridSize=(8,8))enhanced=clahe.apply(denoised). The original data is then binarized to obtain a binary image, which further separates the text from the background. The Otsu algorithm is used to automatically determine the binarization threshold. If the original data is video information rather than image information, the first frame of the video information is used as the input original data.

[0081] Step 402: Set the scanning window. The length and width of the scanning window are both 10% of the original data. Use the scanning window to scan the binary image line by line based on TesseractOCRAPI to extract the text data from the binary image.

[0082] Step 403: Establish a semantic similarity model. The semantic similarity model is based on paraphrase-multilingual-MiniLM-L12-v2 in the sentence-transformers library. It is used to calculate the similarity between text vectors and to match misspelled words with similar shapes or meanings. Input the text data into the semantic similarity model for error correction. After the error correction is completed, the scanned data is obtained.

[0083] The specific steps involved in marking the scale are as follows:

[0084] Step 501: The ruler module counts the pixel size data of the label data, such as the number of pixels occupied by the length and width of different doors and windows in the label data. The average of all pixel size data is used to obtain the pixel ratio, which is the ratio of the actual size of a single pixel.

[0085] Step 502: Set the scale to 20 pixels, multiply the actual size corresponding to the pixel ratio by 20 and assign the value to the scale, and mark it in the second data.

[0086] The verification module verifies the floor plan, specifically including the following steps:

[0087] Step 601: The verification module is a manual data backend. After receiving the floor plan, the verification module synchronously retrieves the corresponding raw data from the database.

[0088] Step 602: The manual data backend identifies the differences between the original data and the floor plan to determine whether the floor plan matches the original data;

[0089] Step 603: If the original data matches the floor plan, it is marked as verified. If the original data does not match the floor plan, the manual data editor will mark the location or area on the floor plan that does not match the original data as feedback information.

[0090] The image analysis process is executed in the following steps:

[0091] Step 202: Convert the image information into a grayscale image, gray=cv2.cvtColor(image,cv2.COLOR_BGR2GRAY), and perform Gaussian blur to remove noise from the grayscale image, blurred=cv2.GaussianBlur(gray,(5,5),0), with the kernel size adjusted according to the noise in the image information;

[0092] Step 203: Perform Canny operator edge detection on the denoised grayscale image. Set dual thresholds to control edge sensitivity and distinguish grayscale images. The low threshold is used to retain weak edges, and the high threshold is used to identify strong edges. The ratio of the low threshold area to the high threshold area is controlled to be 1:3. The specific values ​​of the dual thresholds are dynamically adjusted according to the ratio of the low threshold area to the high threshold area. canny_edges=cv2.Canny(blurred,50,150);

[0093] Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image. The horizontal edge command is sobel_x=cv2.Sobel(blurred,cv2.CV_64F,1,0,ksize=3), and the vertical edge command is sobel_y=cv2.Sobel(blurred,cv2.CV_64F,0,1,ksize=3). The edge enhancement command is sobel_edges=np.uint8(np.absolute(np.hstack([sobel_x,sobel_y]))). The weak edges of the Canny operator are merged with the horizontal or vertical edges of the Sobel operator, and the union is used to enhance the weak edges. The broken edges are connected using the first dilation and then erosion image processing method to obtain the total edge map. The broken edges include weak edges, strong edges, horizontal edges, and vertical edges.

[0094] Step 205: Use Hough transform to extract line segments from the total edge map. The extracted line segments are potential wall lines or door frame lines. The extraction command is:

[0095] defextract_lines(edges):

[0096] lines=cv2.HoughLinesP(edges,1,np.pi / 180,threshold=30,

[0097] minLineLength=50,maxLineGap=10)

[0098] returnlinesiflinesisnotNoneelse[ ];

[0099] It should be noted that, since the feature data carried by the image information is limited, steps 202 to 205 are required to preprocess the image information and extract the edge features in the image information. The edge features refer to the straight line segments in the total edge map, thereby improving the accuracy of subsequent key point recognition.

[0100] Step 206: Establish the first key point model. The image information retains the RGB three channels as the basic features and is input into the first key point model. Then, the extracted line segments are used as auxiliary features and input into the first key point model. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features. After training, the key points in the image information are marked. The key points are the intersection points of edge line segments in the floor plan, such as wall corners or door and window frame corners.

[0101] Step 207: Cross-validate the overlap between the line segments in Step 205 and the key points in Step 206. Set the overlap threshold to 95%. If the overlap rate between the intersection points of the line segments and the key points is greater than or equal to the overlap threshold, it means that the recognition rate of the line segments and key points can meet the requirements for forming the house frame. Combine the line segments and key points to obtain the house frame, and the image analysis process ends. If the overlap rate between the intersection points of the line segments and the key points is less than the overlap threshold, it means that the recognition rate of the line segments and key points cannot meet the requirements for forming the house frame. Jump to Step 203 and repeat Steps 203 to 206 until the overlap rate between the intersection points of the line segments and the key points is greater than or equal to the overlap threshold.

[0102] The echo measurement process specifically includes the following steps:

[0103] To construct a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, and then apply the formula... The slope k of the frequency response curve is calculated in real time. f1 is the frequency at any point on the frequency response curve, A1 is the amplitude at any point on the frequency response curve, f2 is the frequency of the next adjacent point of f1 on the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 on the frequency response curve. When the slope k is greater than or equal to 0.5, it means that the frequency response curve increases sharply at that moment, that is, the mobile terminal receives the echo of the fixed-frequency sound wave hitting the wall. The moment corresponding to the slope k being greater than or equal to 0.5 is marked as the echo point t2. The audio information stores the time t1 when the mobile terminal emits the fixed-frequency sound wave. Subtracting the time t1 from the echo point t2 gives the time difference Δt of the fixed-frequency sound wave's round-trip propagation. According to the formula... The distance d from the moving terminal to the wall in a single direction is obtained, where v is the speed of sound under standard atmospheric pressure.

[0104] like Figure 1As shown, the intelligent floor plan recognition device based on image recognition includes a data acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module, and an output module. The output of the data acquisition module is connected to the input of the database. The output of the database is connected to the inputs of the OpenPose library, the ResNet module, and the OCR module, respectively. The output of the OpenPose library is connected to the input of the ResNet module. The output of the ResNet module is connected to the input of the OCR module. The output of the OCR module is connected to the input of the ruler module. The output of the ruler module is connected to the input of the verification module. The output of the verification module is connected to the input of the output module. The port of the verification module establishes bidirectional communication with the port of the database.

[0105] The acquisition module is used to acquire raw data from the mobile terminal; the OpenPose library is used to identify the floor plan framework of the raw data in the database; the ResNet module is used to identify the identifier data of the raw data; the OCR module is used to extract the scanned data from the raw data; the scale module is used to set the scale; the verification module is used to verify whether the floor plan matches the raw data; the database is used to store the raw data, the feedback information from the verification module, and the computer execution program instructions; and the output module is used to output the floor plan that has passed verification.

[0106] The mobile terminal includes a speaker, an image sensor, and a microphone. The speaker is used to generate fixed-frequency sound waves, the microphone is used to acquire the sound waves that bounce back after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or video information.

[0107] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for intelligent recognition of floor plans based on image recognition, characterized in that: Includes the following steps: Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage. Step 2: The OpenPose library retrieves raw data or feedback information from the database, executes the key point recognition program to obtain the floor plan framework, and then transfers the floor plan framework to the ResNet module. Step 3: The ResNet module retrieves raw data or feedback information from the database, executes the heat map analysis program to obtain the identification data, and merges the identification data and the floor plan framework into the first data and transmits the first data to the OCR module. Step 4: The OCR module retrieves raw data or feedback information from the database, executes the text extraction program to obtain scanned data, and merges the first data and scanned data into second data and transmits the second data to the ruler module. Step 5: The scale module marks the scale in the second data to obtain the floor plan, and the scale module transmits the floor plan to the verification module; Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which then outputs the floor plan for users to view. If the verification fails, the verification module marks the feedback information on the floor plan, uploads the feedback information to the database, and re-executes Step 2.

2. The intelligent floor plan recognition method based on image recognition according to claim 1, characterized in that, The key point recognition program specifically includes the following steps: Step 201: The OpenPose library determines whether the raw data is image information or video information. If it is image information, it executes the image analysis process; if it is video information, it executes the video analysis process. The video analytics process executes in the following steps: Step 208: Separate the picture information and audio information from the video information. The picture information consists of consecutive frames in time sequence. Each consecutive frame is consistent with the image information. Calculate the picture change rate between adjacent consecutive frames pixel by pixel. Set the consecutive frames with a picture change rate greater than or equal to 30% as typical frames. Step 209: Establish the second key point model. The typical frame retains the RGB three channels as continuous features and inputs them into the second key point model. The first key point model is based on the VGG19 convolutional neural network. The second key point model is trained by row convolution on the continuous features. After training, the continuous key points and continuous edge features in the typical frame are marked. In this process, when the mobile terminal is capturing video information, it uses its own speaker to play a fixed-frequency sound wave. The fixed-frequency sound wave spreads outward, hits the wall, and bounces back to the mobile terminal, where the audio information is then obtained by the mobile terminal's microphone. Step 210: Execute the echo measurement process to obtain the distance d from the mobile terminal to the wall in a single direction; Step 211: Repeat the echo measurement process to calculate the distance d in multiple directions, and then combine and stitch the distances d in multiple directions to obtain several audio spaces; Step 212: Cross-validate the continuous keypoints and continuous edge features from Step 209 with several audio spaces from Step 211. Mark the audio space with the highest overlap rate with the continuous keypoints and continuous edge features as audio features, delete the remaining audio spaces, and correct the continuous keypoints and continuous edge features based on the audio features to obtain the house frame.

3. The intelligent floor plan recognition method based on image recognition according to claim 1, characterized in that, The heat map analysis program specifically includes the following steps: Step 301: The ResNet module adjusts the pixel size of the original data and normalizes the pixel values ​​of the original data; Step 302: Load the pre-trained ResNet model, fine-tune it based on the apartment type identification dataset, and use an activation function in the output layer to make the pixel value range of each channel [0,1], thus obtaining a multi-channel heatmap; Step 303: Set a confidence threshold T for the heatmap of each channel, retain regions with pixel values ​​greater than the confidence threshold T, and obtain a binarized mask. The filtering formula is as follows: In the filtering formula, Mc is the mask of the identifier, which determines whether the identifier is retained at the pixel coordinates (x, y), where (x, y) are the pixel coordinates, and Hc(x, y) is the heat map of the identifier. Step 304: Match the heat map of each channel after filtering in step 303 with the house layout framework for location verification, and delete the heat map that fails the location verification. Step 305: Summarize the heat maps verified in Step 304 into identification data.

4. The intelligent floor plan recognition method based on image recognition according to claim 1, characterized in that, The text extraction procedure specifically includes the following steps: Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the original data, improves the contrast between text and background through adaptive histogram equalization, and then binarizes the original data to obtain a binary image. The Otsu algorithm is used to automatically determine the binarization threshold. Step 402: Set up a scanning window, use the scanning window based on TesseractOCRAPI to scan the binary image line by line, and extract the text from the binary image to obtain text data; Step 403: Establish a semantic similarity model, input the text data into the semantic similarity model for error correction, and obtain the scanned data after the error correction is completed.

5. The intelligent floor plan recognition method based on image recognition according to claim 1, characterized in that, Marking a scale involves the following steps: Step 501: The ruler module counts the pixel size data of the identification data, and averages all the pixel size data to obtain the pixel ratio, which is the ratio of the actual size of a single pixel; Step 502: Set the scale, assign the actual dimensions to the scale, and mark them in the second data.

6. The intelligent floor plan recognition method based on image recognition according to claim 1, characterized in that, The verification module verifies the floor plan specifically through the following steps: Step 601: The verification module is a manual data backend. After receiving the floor plan, the verification module synchronously retrieves the corresponding raw data from the database. Step 602: The manual data backend identifies the differences between the original data and the floor plan to determine whether the floor plan matches the original data; Step 603: If the original data matches the floor plan, it is marked as verified. If the original data does not match the floor plan, the manual data editor will mark the location or area on the floor plan that does not match the original data as feedback information.

7. The intelligent floor plan recognition method based on image recognition according to claim 2, characterized in that, The image analysis process executes the following steps: Step 202: Convert the image information into a grayscale image, and then perform Gaussian blur to remove noise from the grayscale image; Step 203: Perform Canny operator edge detection on the grayscale image, and set dual thresholds to control the edge sensitivity to distinguish grayscale images. The low threshold is used to retain weak edges, and the high threshold is used to identify strong edges. Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image, merge the weak edges of the Canny operator with the horizontal or vertical edges of the Sobel operator, take the union, and use the image processing method of dilation followed by erosion to connect the broken edges to obtain the total edge map. Step 205: Use Hough transform to extract line segments from the total edge map; Step 206: Establish the first key point model. The image information retains the RGB three channels as the basic features and inputs them into the first key point model. Then, the extracted line segments are used as auxiliary features and input into the first key point model. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features. After training, the key points in the image information are marked. Step 207: Cross-validate the overlap between the line segments in Step 205 and the key points in Step 206, and set an overlap threshold. If the overlap rate between the intersection points of the line segments and the key points is greater than or equal to the overlap threshold, combine the line segments and key points to obtain the house frame, and the image analysis process ends. If the overlap rate between the intersection points of the line segments and the key points is less than the overlap threshold, jump to Step 203 to execute.

8. The intelligent floor plan recognition method based on image recognition according to claim 2, characterized in that, The echo measurement process specifically includes the following steps: To construct a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, and then apply the formula... The slope k of the frequency response curve is calculated in real time. f1 is the frequency of any point on the frequency response curve, A1 is the amplitude of any point on the frequency response curve, f2 is the frequency of the next adjacent point of f1 on the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 on the frequency response curve. When the slope k is greater than or equal to 0.5, the time corresponding to the slope k being greater than or equal to 0.5 is marked as the echo point t2. The audio information stores the time t1 when the mobile terminal emits a fixed-frequency sound wave. Subtracting time t1 from the echo point t2 gives the time difference Δt of the fixed-frequency sound wave's round-trip propagation. According to the formula... The distance d from the moving terminal to the wall in a single direction is obtained, where v is the speed of sound under standard atmospheric pressure.

9. An intelligent floor plan recognition device based on image recognition, characterized in that... The method for intelligent recognition of floor plans based on image recognition as described in any one of claims 1-8 includes a data acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module, and an output module. The output of the data acquisition module is connected to the input of the database. The output of the database is connected to the inputs of the OpenPose library, the ResNet module, and the OCR module, respectively. The output of the OpenPose library is connected to the input of the ResNet module. The output of the ResNet module is connected to the input of the OCR module. The output of the OCR module is connected to the input of the ruler module. The output of the ruler module is connected to the input of the verification module. The output of the verification module is connected to the input of the output module. The port of the verification module establishes bidirectional communication with the port of the database. The acquisition module is used to acquire raw data from the mobile terminal; the OpenPose library is used to identify the floor plan framework of the raw data in the database; the ResNet module is used to identify the identifier data of the raw data; the OCR module is used to extract the scanned data from the raw data; the scale module is used to set the scale; the verification module is used to verify whether the floor plan matches the raw data; the database is used to store the raw data, the feedback information from the verification module, and the computer execution program instructions; and the output module is used to output the floor plan that has passed verification.

10. The intelligent floor plan recognition device based on image recognition according to claim 9, characterized in that, The mobile terminal includes a speaker, an image sensor, and a microphone. The speaker is used to generate fixed-frequency sound waves, the microphone is used to acquire the sound waves that bounce back after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or video information.

Citation Information

Patent Citations

  • House type image generation method and device based on indoor image, equipment and storage medium

    CN110189398A

  • Human body key point depth relation predication method, device medium and equipment

    CN108830139A

  • Point cloud splicing method based on apartment plan assistance, laser radar and system

    CN116630150A