Intelligent house type image recognition method and device based on image recognition

Through the OpenPose library, ResNet module, OCR module and fixed-frequency acoustic echo measurement technology, the problems of complex structures and low-quality images in floor plan recognition are solved, accurate measurement of floor space and text recognition are achieved, and the integrity and reliability of floor plans are improved.

CN120726641AActive Publication Date: 2025-09-30BBMG TIANTAN FURNITURE CO LTD +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511189834.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2025-09-30
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

Existing intelligent floor plan recognition technology lacks accuracy when identifying complex structures, has difficulty processing low-quality images, lacks the ability to process video information, cannot accurately measure space dimensions, and has low text recognition accuracy, affecting the integrity and usability of floor plans.

Method used

The OpenPose library is used for key point recognition, combined with the ResNet module for heat map analysis, the OCR module is used for text extraction, and the ruler module is used to calculate the scale. Fixed-frequency acoustic echo measurement technology is used for spatial measurement, and a manual verification mechanism is introduced to ensure accuracy.

Benefits of technology

It significantly improves the recognition accuracy of key structures in floor plans, enhances the ability to process low-quality images, achieves accurate measurement of floor space, improves text recognition accuracy, ensures the quality and reliability of floor plans, and realizes the automatic conversion from raw data to standard floor plans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726641A_ABST
    Figure CN120726641A_ABST
Patent Text Reader

Abstract

The invention discloses a house type image intelligent identification method and device based on image identification, and relates to the technical field of house type image identification. The method comprises the following steps that an acquisition module acquires original data from a mobile terminal, the original data comprises image information or video information, the acquisition module forwards the original data to a database for storage, an OpenPose library acquires the original data or feedback information from the database, a key point recognition program is executed to obtain a house type framework, and the house type framework is stored in the database. The OpenPose library transmits the house type framework to the ResNet module, the ResNet module obtains original data or feedback information from the database and executes a hotspot map analysis program to obtain identification data, and the ResNet module combines the identification data and the house type framework into first data and transmits the first data to the OCR module. The method improves the recognition accuracy and efficiency, and solves the technical problems of low accuracy and poor efficiency of a conventional house type image recognition method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of floor plan recognition, and in particular to a floor plan intelligent recognition method and device based on image recognition. Background Art

[0002] In the field of floor plan generation, existing technologies are primarily categorized into traditional manual drawing and semi-automatic recognition based on image recognition. Traditional methods rely on on-site measurements (such as using a tape measure or laser rangefinder) and manual drawing using CAD software. Data such as wall lengths and door and window locations are manually recorded and converted into graphical representations. With the advancement of computer vision technology, auxiliary methods based on image recognition have emerged. These utilize algorithms such as edge detection and feature extraction to identify elements such as walls, doors, and windows from house images. Some methods integrate optical character recognition (OCR) technology to extract room names, assisting in the initial drawing of floor plans.

[0003] Patent publication number CN110189398A discloses a method for generating floor plans based on indoor images. The method uses a preset image recognition model to identify the initial image, determine whether it is an indoor image, and detect the floor plan generation components in the indoor image. The three-dimensional floor plan is generated based on the floor plan generation components and their position and size information. However, the existing floor plan intelligent recognition technology still has some problems and shortcomings: the existing technology is not accurate enough in identifying key structures such as walls, doors and windows in the floor plan. In particular, for floor plans with complex structures, the recognition accuracy is low and it is difficult to meet the actual application requirements. For floor plans with poor image quality, For blurred, deformed or poorly lit images, the recognition accuracy of existing technologies is significantly reduced, and they are unable to effectively process floor plans under various quality conditions. Existing technologies mainly process static images and lack the ability to effectively extract and process floor plan data in video information, which limits the application scenarios of the technology. Existing floor plan recognition methods lack accurate measurement methods for the actual dimensions of floor plans, making it difficult to provide accurate spatial dimension information, affecting subsequent decoration design and space planning. Existing technologies have deficiencies in text information extraction, especially for handwritten annotations or special fonts, where the recognition accuracy is low, affecting the integrity and usability of floor plans.

[0004] Therefore, there is an urgent need for an intelligent floor plan recognition method that can accurately identify key structures in floor plans, adapt to different image qualities, process video information, accurately measure space dimensions and accurately extract text information to solve existing technical problems. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a method and device for intelligently identifying floor plans based on image recognition, which solves the problems raised in the above-mentioned background technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a method for intelligently identifying floor plans based on image recognition, comprising the following steps: Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage; Step 2: The OpenPose library obtains raw data or feedback information from the database, executes the key point recognition program to obtain the house frame, and transmits the house frame to the ResNet module; Step 3: The ResNet module obtains raw data or feedback information from the database, executes the heat map analysis program to obtain identification data, and the ResNet module combines the identification data and the house type framework into first data and transmits the first data to the OCR module; Step 4: The OCR module obtains the original data or feedback information from the database, executes the text extraction program to obtain the scanned data, and the OCR module merges the first data and the scanned data into the second data and transmits the second data to the ruler module; Step 5: The ruler module marks a scale in the second data to obtain a floor plan, and the ruler module transmits the floor plan to the verification module; Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which outputs the floor plan to a mobile terminal or display platform for users to view. If the verification fails, the verification module marks the feedback information on the floor plan, uploads the feedback information to the database, and re-executes step 2.

[0007] Furthermore, the key point identification procedure specifically includes the following steps: Step 201: The OpenPose library determines whether the original data is image information or video information. If it is image information, the image analysis process is executed; if it is video information, the video analysis process is executed; The video analysis process executes the following steps: Step 208: Separate the image information and audio information from the video information. The image information consists of time-sequential consecutive frames, each of which is consistent with the image information. Since the image information is larger than the original image information data and has sufficient feature data for analysis, several typical frames are extracted from the image information to directly perform key point analysis. The image change rate between adjacent consecutive frames is calculated pixel by pixel, which is the percentage of the number of pixel values ​​that have changed relative to the total number of pixels. Continuous frames with an image change rate greater than or equal to 30% are set as typical frames. Step 209: Establish a second key point model, retain the RGB three channels of the typical frame as continuous features and input them into the second key point model. The first key point model is based on the VGG19 convolutional neural network, and the second key point model performs row convolution training on the continuous features. If there is feedback information in the database and the convolution training output is the same as the feedback information, the convolution training is repeated. After the training is completed, the continuous key points and continuous edge features in the typical frame are marked. When a mobile terminal captures video, it uses its own speakers to play fixed-frequency sound waves in the 500Hz to 2kHz range. The fixed-frequency sound waves diffuse outward, hit the wall, and then bounce back to the mobile terminal. The mobile terminal's microphone then captures the audio information. Essentially, the mobile terminal uses the diffuse echo of the fixed-frequency sound waves to analyze the interior space of the apartment. Step 210: Execute an echo measurement process to obtain a distance d from the mobile terminal to the wall in a single direction; Step 211: Repeat the echo measurement process to calculate the distance d in multiple directions. Because the mobile terminal uses a single microphone, the collected audio information is not directional, and the specific direction of the distance d cannot be determined. The distances d in multiple directions are arranged and combined to obtain several audio spaces. At this time, the spatial dimensions cannot determine the specific framework, and multiple possible results are output. These are then verified by the continuous key points and continuous edge features output by the subsequent second key point model. Step 212: Cross-validate the overlap of the continuous key points and continuous edge features in step 209 with several audio spaces in step 211, mark the audio space with the highest overlap rate with the continuous key points and continuous edge features as the audio feature, delete the remaining audio spaces, and use the audio features as the basis. The audio features are calculated based on fixed-frequency sound waves, and the calculation results about the wall distance are more accurate and reliable. The continuous key points and continuous edge features are corrected to obtain the apartment framework, that is, when there is an error in the distance between the continuous key points and continuous edge features and the audio features, the distance of the audio features shall prevail.

[0008] Furthermore, the heat map analysis program specifically includes the following steps: Step 301: The ResNet module adjusts the pixel size of the original data and normalizes the pixel values ​​of the original data. If the original data is video information rather than image information, the first frame of the picture information in the video information is used as the input original data; Step 302: Load a pre-trained ResNet model, such as ResNet-50, and fine-tune it based on a household identification dataset. The household identification dataset includes labeled data for single doors, double doors, sliding doors, windows, etc., which are manually labeled in advance by data annotators. The household identification dataset is stored in a database. The output layer uses a Sigmoid activation function to set the pixel value range of each channel to [0, 1], thus obtaining a multi-channel heat map. If feedback information exists in the database and the heat map output result is the same as the feedback information, the heat map is re-output. Step 303: Set the confidence threshold T for each channel’s heat map, for example, T=0.6, retain the area where the pixel value is greater than the confidence threshold T, filter the low probability noise, and obtain a binary mask. The filtering formula is: , where Mc is the mask of the marker, which determines whether the marker is retained at the pixel coordinate (x, y), (x, y) is the pixel coordinate, and Hc (x, y) is the heat map of the marker. The original heat map contains noise and redundant candidate areas. Post-processing can filter out marker areas with high confidence and clear boundaries, thereby improving the accuracy of subsequent recognition; Step 304: Match the heat map of each channel after filtering in step 303 with the apartment framework for position verification. The apartment framework is the wall direction output by the OpenPose library. Delete the heat map that fails the position verification to ensure that the identification is consistent with the spatial logic of the apartment framework, providing reliable information for the subsequent merging of the first data. Step 305: The heat map verified in step 304 is aggregated into identification data.

[0009] Furthermore, the text extraction program specifically includes the following steps: Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the raw data, improves the contrast between the text and the background through adaptive histogram equalization, and then binarizes the raw data to obtain a binary image to further separate the text and the background. The Otsu algorithm is used to automatically determine the binarization threshold. If the raw data is video information rather than image information, the first frame of the image information in the video information is used as the input raw data; Step 402: Set a scanning window, where the length and width of the scanning window are both 10% of the original data. Use the scanning window to scan the binary image line by line based on the TesseractOCRAPI to extract the text in the binary image to obtain text data. If there is feedback information in the database and the output text data is the same as the feedback information, re-execute step 402 to output the text data. Step 403: Establish a semantic similarity model. The semantic similarity model is based on paraphrase-multilingual-MiniLM-L12-v2 in the sentence-transformers library and is used to calculate the similarity between text vectors. It is used to match misspelled characters with similar glyphs or semantics. The text data is input into the semantic similarity model for error correction. After the error correction is completed, the scanned data is obtained.

[0010] Furthermore, marking the scale specifically includes the following steps: Step 501: The ruler module counts the pixel size data of the identification data, for example, the number of pixels occupied by the length and width of different doors and windows in the identification data, and averages all the pixel size data to obtain a pixel ratio, which is the ratio of a single pixel to the actual size. Step 502: Set a scale, multiply the actual size corresponding to the pixel ratio by 20, assign the value to the scale, and mark it in the second data.

[0011] Furthermore, the verification module verifies the floor plan and specifically includes the following steps: Step 601: The verification module is specifically a manual data background. After receiving the floor plan, the verification module synchronously obtains the corresponding original data from the database; Step 602: The manual data backend determines whether the floor plan matches the original data by identifying the differences between the original data and the floor plan; Step 603: If the original data matches the floor plan, it is marked as passed. If the original data does not match the floor plan, the manual data will mark the location or area in the floor plan that does not match the original data as feedback information.

[0012] Furthermore, the image analysis process performs the following steps: Step 202: converting the image information into a grayscale image, performing Gaussian blurring on the grayscale image to remove noise, and adjusting the function kernel size according to the noise of the image information; Step 203: Perform Canny operator edge detection on the de-noised grayscale image, set a dual threshold to control edge sensitivity to distinguish the grayscale image, the low threshold is used to retain weak edges, and the high threshold is used to determine strong edges. The ratio of the low threshold area to the high threshold area is controlled to be 1:3, and the specific value of the dual threshold is dynamically adjusted according to the ratio of the low threshold area to the high threshold area; Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image. Merge the weak edges of the Canny operator with the horizontal or vertical edges of the Sobel operator, and take the union to enhance the weak edges. Use the image processing method of first dilation and then erosion to connect the broken edges to obtain a total edge map. The broken edges include weak edges, strong edges, horizontal edges, and vertical edges. Step 205: using Hough transform to extract straight line segments in the total edge map, where the extracted straight line segments are potential wall straight lines or door frame straight lines; It should be noted that, since the feature data carried by the image information is limited, steps 202 to 205 are required to pre-process the image information to extract edge features in the image information. Edge features refer to straight line segments in the total edge map, thereby improving the accuracy of subsequent key point recognition. Step 206: Establish a first key point model, retain the RGB three channels of the image information as basic features and input them into the first key point model, then input the extracted straight line segments as auxiliary features into the first key point model. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features. After the training is completed, the key points in the image information are marked. The key points are the intersection points of the edge segments in the floor plan. Step 207: Cross-validate the line segments in step 205 by overlapping with the key points in step 206, and set an overlap threshold. If the overlap rate between the intersection points between the line segments and the key points is greater than or equal to the overlap threshold, it means that the recognition rate of the line segments and the key points can meet the requirements for forming the apartment frame. The line segments and the key points are combined to obtain the apartment frame, and the image analysis process ends. If the overlap rate between the intersection points between the line segments and the key points is less than the overlap threshold, it means that the recognition rate of the line segments and the key points cannot meet the requirements for forming the apartment frame. Jump to step 203 to execute, and repeat steps 203 to 206 until the overlap rate between the intersection points between the line segments and the key points is greater than or equal to the overlap threshold.

[0013] Furthermore, the echo measurement process specifically includes the following steps: Create a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, according to the formula The slope k of the frequency response curve is calculated in real time. f1 is the frequency of any point in the frequency response curve, A1 is the amplitude of any point in the frequency response curve, f2 is the frequency of the next adjacent point of f1 in the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 in the frequency response curve. When the slope k is greater than or equal to 0.5, it means that the frequency response curve at that moment increases sharply, that is, the mobile terminal receives the echo of the fixed-frequency sound wave hitting the wall. The moment corresponding to the slope k greater than or equal to 0.5 is marked as the echo point t2. The time t1 when the mobile terminal sends the fixed-frequency sound wave is saved in the audio information. The time difference Δt of the fixed-frequency sound wave propagating back and forth is obtained by subtracting the time t1 from the echo point t2. According to the formula The distance d from the moving terminal to the wall in a single direction is obtained, and v is the speed of sound under standard atmospheric pressure.

[0014] An intelligent floor plan recognition device based on image recognition includes an acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module, and an output module. The output end of the acquisition module is connected to the input end of the database, the output end of the database is respectively connected to the input ends of the OpenPose library, the ResNet module, and the OCR module, the output end of the OpenPose library is connected to the input end of the ResNet module, the output end of the ResNet module is connected to the input end of the OCR module, the output end of the OCR module is connected to the input end of the ruler module, the output end of the ruler module is connected to the input end of the verification module, the output end of the verification module is connected to the input end of the output module, and the port of the verification module establishes two-way communication with the port of the database; The acquisition module is used to obtain raw data from the mobile terminal, the OpenPose library is used to identify the floor plan framework of the raw data in the database, the ResNet module is used to identify the identification data of the raw data, the OCR module is used to extract the scan data from the raw data, the ruler module is used to set the scale, the verification module is used to verify whether the floor plan matches the raw data, the database is used to store the raw data, feedback information from the verification module, and computer execution program instructions, and the output module is used to output the verified floor plan to the outside.

[0015] Furthermore, the mobile terminal includes a speaker, an image sensor and a microphone. The speaker is used to generate fixed-frequency sound waves to the outside, the microphone is used to obtain the sound waves rebounding after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or picture information in the video information.

[0016] The beneficial effects of the present invention are as follows: by adopting the OpenPose library to identify key points and distinguish image and video processing processes, and combining multiple image processing technologies such as grayscale image conversion, Gaussian blur denoising, Canny operator edge detection, etc., the accuracy of floor plan recognition is significantly improved, and key structures such as walls, doors and windows can be accurately identified, while the processing capability of low-quality images is enhanced; the fixed-frequency acoustic wave echo measurement technology is innovatively applied to video information processing, which realizes the accurate measurement of the actual size of the floor plan space and expands the application scenarios; the ResNet model is used to perform heat map analysis, which improves the recognition accuracy of floor plan identification elements such as doors and windows; the OCR module combines the semantic similarity model to improve the accuracy of text recognition; the ruler module establishes a scale by calculating the pixel ratio, which realizes the accurate measurement of the actual size of the floor plan space; the manual verification mechanism ensures the quality and reliability of the floor plan; the overall automatic conversion from raw data to standard floor plan is realized, which improves work efficiency and has obvious advantages over the existing technology.

[0017] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 This is a system block diagram of the intelligent floor plan recognition device based on image recognition of the present invention. DETAILED DESCRIPTION

[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0021] See also Figure 1 The present invention provides a technical solution: a method for intelligently identifying floor plans based on image recognition, comprising the following steps: Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage; Step 2: The OpenPose library obtains raw data or feedback information from the database, executes the key point recognition program to obtain the house frame, and transmits the house frame to the ResNet module; Step 3: The ResNet module obtains raw data or feedback information from the database, executes the heat map analysis program to obtain identification data, and the ResNet module combines the identification data and the house type framework into first data and transmits the first data to the OCR module; Step 4: The OCR module obtains the original data or feedback information from the database, executes the text extraction program to obtain the scanned data, and the OCR module merges the first data and the scanned data into the second data and transmits the second data to the ruler module; Step 5: The ruler module marks a scale in the second data to obtain a floor plan, and the ruler module transmits the floor plan to the verification module; Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which outputs the floor plan to a mobile terminal or display platform for users to view. If the verification fails, the verification module marks the feedback information on the floor plan, uploads the feedback information to the database, and re-executes step 2.

[0022] The key point identification procedure specifically includes the following steps: Step 201: The OpenPose library determines whether the original data is image information or video information. If it is image information, the image analysis process is executed; if it is video information, the video analysis process is executed; The video analysis process executes the following steps: Step 208: Separate the image information and audio information from the video information. The image information consists of time-sequential consecutive frames, each of which is consistent with the image information. Since the image information is larger than the original image information data and has sufficient feature data for analysis, several typical frames are extracted from the image information to directly perform key point analysis. The image change rate between adjacent consecutive frames is calculated pixel by pixel, which is the percentage of the number of pixel values ​​that have changed relative to the total number of pixels. Continuous frames with an image change rate greater than or equal to 30% are set as typical frames. Step 209: Establish a second key point model, retain the RGB three channels of the typical frame as continuous features and input them into the second key point model. The first key point model is based on the VGG19 convolutional neural network, and the second key point model performs row convolution training on the continuous features. After the training is completed, the continuous key points and continuous edge features in the typical frame are marked; When a mobile terminal captures video, it uses its own speakers to play fixed-frequency sound waves in the 500Hz to 2kHz range. The fixed-frequency sound waves diffuse outward, hit the wall, and then bounce back to the mobile terminal. The mobile terminal's microphone then captures the audio information. Essentially, the mobile terminal uses the diffuse echo of the fixed-frequency sound waves to analyze the interior space of the apartment. Step 210: Execute an echo measurement process to obtain a distance d from the mobile terminal to the wall in a single direction; Step 211: Repeat the echo measurement process to calculate the distance d in multiple directions. Because the mobile terminal uses a single microphone, the collected audio information is not directional, and the specific direction of the distance d cannot be determined. The distances d in multiple directions are arranged and combined to obtain several audio spaces. At this time, the spatial dimensions cannot determine the specific framework, and multiple possible results are output. These are then verified by the continuous key points and continuous edge features output by the subsequent second key point model. Step 212: Cross-validate the overlap of the continuous key points and continuous edge features in step 209 with several audio spaces in step 211, mark the audio space with the highest overlap rate with the continuous key points and continuous edge features as the audio feature, delete the remaining audio spaces, and use the audio features as the basis. The audio features are calculated based on fixed-frequency sound waves, and the calculation results about the wall distance are more accurate and reliable. The continuous key points and continuous edge features are corrected to obtain the apartment framework, that is, when there is an error in the distance between the continuous key points and continuous edge features and the audio features, the distance of the audio features shall prevail.

[0023] The heat map analysis procedure specifically includes the following steps: Step 301: The ResNet module resizes the original data to a uniform size of 224×224 pixels to match the default input of ResNet-50. The module also normalizes the pixel values ​​of the original data, converting them proportionally from the range [0, 255] to [0, 1]. If the original data is video information rather than image information, the first frame of the video information is used as the input raw data. Step 302: Load a pre-trained ResNet model, such as ResNet-50, and fine-tune it based on a household identification dataset. The household identification dataset includes labeled data for single doors, double doors, sliding doors, windows, etc., which are manually labeled in advance by data annotators. The household identification dataset is stored in a database. The output layer is activated by a Sigmoid function to set the pixel value range of each channel to [0, 1], thus obtaining a multi-channel heat map. Step 303: Set the confidence threshold T for each channel’s heat map, for example, T=0.6, retain the area where the pixel value is greater than the confidence threshold T, filter the low probability noise, and obtain a binary mask. The filtering formula is: , where Mc is the mask of the marker, which determines whether the marker is retained at the pixel coordinate (x, y), (x, y) is the pixel coordinate, and Hc (x, y) is the heat map of the marker. The original heat map contains noise and redundant candidate areas. Post-processing can filter out marker areas with high confidence and clear boundaries, thereby improving the accuracy of subsequent recognition; Step 304: The heat map of each channel filtered in step 303 is matched with the apartment framework for position verification. The apartment framework is the wall orientation output by the OpenPose library. For example, doors and windows must be located on the wall. Heat maps that fail the position verification are deleted to ensure that the identification is consistent with the spatial logic of the apartment framework, providing reliable information for the subsequent merging of the first data. Step 305: The heat map verified in step 304 is aggregated into identification data.

[0024] The text extraction procedure specifically includes the following steps: Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the original data. The instruction is blurred=GaussianBlur(gray,(5,5),0). The contrast between the text and the background is improved through adaptive histogram equalization. The instruction is clahe=CLAHE(clipLimit=2.0,tileGridSize=(8,8))enhanced=clahe.apply(denoised). The original data is then binarized to obtain a binary image to further separate the text and the background. The Otsu algorithm is used to automatically determine the binarization threshold. If the original data is video information rather than image information, the first frame of the picture information in the video information is used as the input original data. Step 402: Set a scanning window, where the length and width of the scanning window are both 10% of the original data, and use the scanning window to scan the binary image line by line based on the TesseractOCRAPI to extract the text in the binary image to obtain text data; Step 403: Establish a semantic similarity model. The semantic similarity model is based on paraphrase-multilingual-MiniLM-L12-v2 in the sentence-transformers library and is used to calculate the similarity between text vectors. It is used to match misspelled characters with similar glyphs or semantics. The text data is input into the semantic similarity model for error correction. After the error correction is completed, the scanned data is obtained.

[0025] The scale marking process specifically includes the following steps: Step 501: The ruler module counts the pixel size data of the identification data, for example, the number of pixels occupied by the length and width of different doors and windows in the identification data, and averages all the pixel size data to obtain a pixel ratio, which is the ratio of a single pixel to the actual size. Step 502: Set the scale to 20 pixels, multiply the actual size corresponding to the pixel ratio by 20, assign the value to the scale and mark it in the second data.

[0026] The verification module verifies the floor plan and specifically includes the following steps: Step 601: The verification module is specifically a manual data background. After receiving the floor plan, the verification module synchronously obtains the corresponding original data from the database; Step 602: The manual data backend determines whether the floor plan matches the original data by identifying the differences between the original data and the floor plan; Step 603: If the original data matches the floor plan, it is marked as passed. If the original data does not match the floor plan, the manual data will mark the location or area in the floor plan that does not match the original data as feedback information.

[0027] The image analysis process executes the following steps: Step 202: Convert the image information into a grayscale image, gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY), and perform Gaussian blurring on the grayscale image to remove noise, blurred = cv2.GaussianBlur(gray, (5, 5), 0), where the kernel size of the function is adjusted according to the noise of the image information; Step 203: Perform Canny operator edge detection on the denoised grayscale image. Set a dual threshold to control edge sensitivity to distinguish the grayscale image. The low threshold is used to retain weak edges, and the high threshold is used to determine strong edges. The ratio of the low threshold area to the high threshold area is controlled to be 1:3. The specific value of the dual threshold is dynamically adjusted according to the ratio of the low threshold area to the high threshold area. canny_edges = cv2.Canny(blurred,50,150); Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image. The horizontal edge instruction is sobel_x=cv2.Sobel(blurred,cv2.CV_64F,1,0,ksize=3), the vertical edge instruction is sobel_y=cv2.Sobel(blurred,cv2.CV_64F,0,1,ksize=3), and the edge enhancement instruction is sobel_edges=np.uint8(np.absolute(np.hstack([sobel_x,sobel_y]))). The weak edge of the Canny operator is fused with the horizontal or vertical edge of the Sobel operator, and the union is taken to enhance the weak edge. The broken edges are connected using an image processing method of first dilation and then erosion to obtain a total edge map. The broken edges include weak edges, strong edges, horizontal edges, and vertical edges. Step 205: Use Hough transform to extract straight line segments in the total edge map. The extracted straight line segments are potential wall lines or door frame lines. The extraction command is: defextract_lines(edges): lines=cv2.HoughLinesP(edges,1,np.pi / 180,threshold=30, minLineLength=50,maxLineGap=10) returnlinesiflinesisnotNoneelse[ ]; It should be noted that, since the feature data carried by the image information is limited, steps 202 to 205 are required to pre-process the image information to extract edge features in the image information. Edge features refer to straight line segments in the total edge map, thereby improving the accuracy of subsequent key point recognition. Step 206: Establish a first key point model. The RGB three channels of the image information are retained as basic features and input into the first key point model. The extracted straight line segments are then input into the first key point model as auxiliary features. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features. After the training is completed, the key points in the image information are marked. The key points are the intersection points of the edge segments in the floor plan, such as the corners of walls or door and window frames. Step 207: Cross-validate the line segments in step 205 by overlapping with the key points in step 206, and set the overlap threshold to 95%. If the overlap rate between the intersection points between the line segments and the key points is greater than or equal to the overlap threshold, it means that the recognition rate of the line segments and the key points can meet the requirements for forming the apartment frame. The line segments and the key points are combined to obtain the apartment frame, and the image analysis process ends. If the overlap rate between the intersection points between the line segments and the key points is less than the overlap threshold, it means that the recognition rate of the line segments and the key points cannot meet the requirements for forming the apartment frame. Jump to step 203 and repeat steps 203 to 206 until the overlap rate between the intersection points between the line segments and the key points is greater than or equal to the overlap threshold.

[0028] The echo measurement process specifically includes the following steps: Create a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, according to the formula The slope k of the frequency response curve is calculated in real time. f1 is the frequency of any point in the frequency response curve, A1 is the amplitude of any point in the frequency response curve, f2 is the frequency of the next adjacent point of f1 in the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 in the frequency response curve. When the slope k is greater than or equal to 0.5, it means that the frequency response curve at that moment increases sharply, that is, the mobile terminal receives the echo of the fixed-frequency sound wave hitting the wall. The moment corresponding to the slope k greater than or equal to 0.5 is marked as the echo point t2. The time t1 when the mobile terminal sends the fixed-frequency sound wave is saved in the audio information. The time difference Δt of the fixed-frequency sound wave propagating back and forth is obtained by subtracting the time t1 from the echo point t2. According to the formula The distance d from the moving terminal to the wall in a single direction is obtained, and v is the speed of sound under standard atmospheric pressure.

[0029] like Figure 1As shown, the intelligent recognition device for floor plan based on image recognition includes an acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module and an output module. The output end of the acquisition module is connected to the input end of the database, the output end of the database is connected to the input ends of the OpenPose library, the ResNet module and the OCR module respectively, the output end of the OpenPose library is connected to the input end of the ResNet module, the output end of the ResNet module is connected to the input end of the OCR module, the output end of the OCR module is connected to the input end of the ruler module, the output end of the ruler module is connected to the input end of the verification module, the output end of the verification module is connected to the input end of the output module, and the port of the verification module establishes two-way communication with the port of the database; The acquisition module is used to obtain the original data of the mobile terminal, the OpenPose library is used to identify the floor plan framework of the original data in the database, the ResNet module is used to identify the identification data of the original data, the OCR module is used to extract the scan data from the original data, the ruler module is used to set the scale, the verification module is used to verify whether the floor plan matches the original data, the database is used to store the original data, the feedback information of the verification module and the computer execution program instructions, and the output module is used to output the verified floor plan to the outside.

[0030] Among them, the mobile terminal includes a speaker, an image sensor and a microphone. The speaker is used to generate fixed-frequency sound waves to the outside, the microphone is used to obtain the sound waves rebounding after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or picture information in video information.

[0031] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An intelligent floor plan recognition method based on image recognition, characterized by: The following steps are involved: Step 1: The acquisition module obtains raw data from the mobile terminal. The raw data includes image information or video information. The acquisition module forwards the raw data to the database for storage; Step 2: The OpenPose library obtains raw data or feedback information from the database, executes the key point recognition program to obtain the house frame, and transmits the house frame to the ResNet module; Step 3: The ResNet module obtains raw data or feedback information from the database, executes the heat map analysis program to obtain identification data, and the ResNet module combines the identification data and the house type framework into first data and transmits the first data to the OCR module; Step 4: The OCR module obtains the original data or feedback information from the database, executes the text extraction program to obtain the scanned data, and the OCR module merges the first data and the scanned data into the second data and transmits the second data to the ruler module; Step 5: The ruler module marks a scale in the second data to obtain a floor plan, and the ruler module transmits the floor plan to the verification module; Step 6: The verification module verifies the floor plan. If the verification passes, the verification module forwards the floor plan to the output module, which outputs the floor plan for the user to view. If the verification fails, the verification module marks the feedback information on the floor plan, uploads the feedback information to the database, and re-executes step 2.

2. The method for intelligently identifying floor plans based on image recognition according to claim 1, characterized in that: The key point identification procedure specifically includes the following steps: Step 201: The OpenPose library determines whether the original data is image information or video information. If it is image information, the image analysis process is executed; if it is video information, the video analysis process is executed; The video analysis process executes the following steps: Step 208: Separate the image information and audio information from the video information. The image information is composed of consecutive frames in time sequence. Each consecutive frame is consistent with the image information. The image change rate between adjacent consecutive frames is calculated pixel by pixel. The consecutive frames with a frame change rate greater than or equal to 30% are set as typical frames. Step 209: Establish a second key point model, retain the RGB three channels of the typical frame as continuous features and input them into the second key point model. The first key point model is based on the VGG19 convolutional neural network, and the second key point model performs row convolution training on the continuous features. After the training is completed, the continuous key points and continuous edge features in the typical frame are marked; When a mobile terminal is shooting video information, it uses its own speaker to play fixed-frequency sound waves. The fixed-frequency sound waves spread outward, hit the wall, and then bounce back to the mobile terminal, and then the mobile terminal's microphone captures the audio information. Step 210: Execute an echo measurement process to obtain a distance d from the mobile terminal to the wall in a single direction; Step 211: Repeat the echo measurement process to calculate the distances d in multiple directions, and arrange and combine the distances d in multiple directions to obtain multiple audio spaces; Step 212: Cross-validate the continuous key point and continuous edge features in step 209 by overlapping them with the audio spaces in step 211. Mark the audio space with the highest overlap rate with the continuous key point and continuous edge features as the audio feature, delete the remaining audio spaces, and modify the continuous key point and continuous edge features based on the audio features to obtain the apartment framework.

3. The method for intelligently identifying floor plans based on image recognition according to claim 1, characterized in that: The heat map analysis procedure specifically includes the following steps: Step 301: The ResNet module adjusts the pixel size of the original data and normalizes the pixel values ​​of the original data; Step 302: Load the pre-trained ResNet model and fine-tune it based on the housing type identification dataset. The output layer uses an activation function to make the pixel value range of each channel [0, 1], thus obtaining a multi-channel heat map. Step 303: Set a confidence threshold T for the heat map of each channel, retain the area with pixel values ​​greater than the confidence threshold T, and obtain a binary mask. The filtering formula is: ,In the filtering formula, Mc is the mask of the marker, which determines whether the marker is retained at the pixel coordinate (x, y), (x, y) is the pixel coordinate, and Hc (x, y) is the heat map of the marker; Step 304: Match the heat map of each channel filtered in step 303 with the apartment frame to perform location verification, and delete the heat map that fails the location verification; Step 305: The heat map verified in step 304 is aggregated into identification data.

4. The method for intelligently identifying floor plans based on image recognition according to claim 1, characterized in that: The text extraction program specifically includes the following steps: Step 401: The OCR module uses a 5x5 convolution kernel Gaussian blur to denoise the original data, improves the contrast between the text and the background through adaptive histogram equalization, and then binarizes the original data to obtain a binary image. The Otsu algorithm is used to automatically determine the binarization threshold. Step 402: Setting a scanning window, scanning the binary image line by line using the scanning window based on TesseractOCRAPI, extracting text from the binary image to obtain text data; Step 403: Establish a semantic similarity model, input the text data into the semantic similarity model for error correction, and obtain the scanned data after the error correction is completed.

5. The method for intelligently identifying floor plans based on image recognition according to claim 1, characterized in that: Marking the scale bar specifically includes the following steps: Step 501: The ruler module counts pixel size data of the identification data, averages all pixel size data to obtain a pixel ratio, where the pixel ratio is the ratio of a single pixel to its actual size; Step 502: Set the scale, assign the actual size to the scale and mark it in the second data.

6. The method for intelligently identifying floor plans based on image recognition according to claim 1, characterized in that: The verification module verifies the floor plan and specifically includes the following steps: Step 601: The verification module is specifically a manual data background. After receiving the floor plan, the verification module synchronously obtains the corresponding original data from the database; Step 602: The manual data backend determines whether the floor plan matches the original data by identifying the differences between the original data and the floor plan; Step 603: If the original data matches the floor plan, it is marked as passed. If the original data does not match the floor plan, the manual data will mark the location or area in the floor plan that does not match the original data as feedback information.

7. The method for intelligently identifying floor plans based on image recognition according to claim 2, characterized in that: The image analysis process performs the following steps: Step 202: converting the image information into a grayscale image, and performing Gaussian blurring on the grayscale image to remove noise; Step 203: Perform Canny operator edge detection on the grayscale image, set dual thresholds to control edge sensitivity to distinguish the grayscale image, the low threshold is used to retain weak edges, and the high threshold is used to determine strong edges; Step 204: Use the Sobel operator to enhance the horizontal or vertical edges of the grayscale image, fuse the weak edges of the Canny operator with the horizontal or vertical edges of the Sobel operator, take the union, and use the image processing method of first dilation and then erosion to connect the broken edges to obtain the total edge map; Step 205: extracting straight line segments in the total edge map using Hough transform; Step 206: Establish a first key point model, retain the RGB three channels of the image information as basic features and input them into the first key point model, then input the extracted straight line segments as auxiliary features into the first key point model. The first key point model is based on the VGG19 convolutional neural network. The first key point model performs convolution training on the basic features and auxiliary features, and after the training is completed, the key points in the image information are marked; Step 207: Cross-validate the overlapped straight line segments in step 205 with the key points in step 206, and set an overlap threshold. If the overlap ratio between the intersection points between the straight line segments and the key points is greater than or equal to the overlap threshold, combine the straight line segments and the key points to obtain the apartment frame, and the image analysis process ends. If the overlap ratio between the intersection points between the straight line segments and the key points is less than the overlap threshold, jump to step 203 for execution.

8. The method for intelligently identifying floor plans based on image recognition according to claim 2, characterized in that: The echo measurement process specifically includes the following steps: Create a frequency response diagram, input the audio information into the frequency response diagram to obtain the frequency response curve of the audio information, according to the formula The slope k of the frequency response curve is calculated in real time. f1 is the frequency of any point in the frequency response curve, A1 is the amplitude of any point in the frequency response curve, f2 is the frequency of the next adjacent point of f1 in the frequency response curve, and A2 is the amplitude of the next adjacent point of A1 in the frequency response curve. When the slope k is greater than or equal to 0.5, the moment corresponding to the slope k being greater than or equal to 0.5 is marked as the echo point t2. The time t1 when the mobile terminal sends the fixed-frequency sound wave is stored in the audio information. The time difference Δt of the fixed-frequency sound wave propagating back and forth is obtained by subtracting the time t1 from the echo point t2. According to the formula The distance d from the moving terminal to the wall in a single direction is obtained, and v is the speed of sound under standard atmospheric pressure.

9. The intelligent recognition device for floor plan based on image recognition is characterized by Implementing the method for intelligent floor plan recognition based on image recognition as described in any one of claims 1 to 8, comprising an acquisition module, a database, an OpenPose library, a ResNet module, an OCR module, a ruler module, a verification module and an output module, wherein the output end of the acquisition module is connected to the input end of the database, the output end of the database is respectively connected to the input ends of the OpenPose library, the ResNet module and the OCR module, the output end of the OpenPose library is connected to the input end of the ResNet module, the output end of the ResNet module is connected to the input end of the OCR module, the output end of the OCR module is connected to the input end of the ruler module, the output end of the ruler module is connected to the input end of the verification module, the output end of the verification module is connected to the input end of the output module, and the port of the verification module establishes bidirectional communication with the port of the database; The acquisition module is used to obtain raw data from the mobile terminal, the OpenPose library is used to identify the floor plan framework of the raw data in the database, the ResNet module is used to identify the identification data of the raw data, the OCR module is used to extract the scan data from the raw data, the ruler module is used to set the scale, the verification module is used to verify whether the floor plan matches the raw data, the database is used to store the raw data, feedback information from the verification module, and computer execution program instructions, and the output module is used to output the verified floor plan to the outside.

10. The intelligent floor plan recognition device based on image recognition according to claim 9, characterized in that: The mobile terminal includes a speaker, an image sensor and a microphone. The speaker is used to generate fixed-frequency sound waves, the microphone is used to obtain the sound waves rebounded after the fixed-frequency sound waves hit the wall, and the image sensor is used to capture image information or picture information in video information.

Citation Information

Patent Citations

  • Human body key point depth relation predication method, device medium and equipment

    CN108830139A

  • Point cloud splicing method based on apartment plan assistance, laser radar and system

    CN116630150A

  • Inertial navigation deviation correction method, device and equipment based on echo detection and medium

    CN116659554A

  • House type identification and three-dimensional reconstruction method and system based on grating image

    CN120279394A

  • Standardized process algorithm for electronic drawing of interior design and for improving ai recognition rate

    WO2025091732A1