Customized cabinet generation method and device, electronic equipment and storage medium
Through the multi-scale feature pyramid network and hierarchical feature alignment strategy, the rapid and accurate generation of customized cabinets is achieved, which solves the problems of low design efficiency and serious homogeneity in home customization and meets the diverse customization needs of users.
Patent Information
- Application Number
- CN202510816020.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-03
AI Technical Summary
Existing technologies in home customization have problems such as low design efficiency, large dimensional deviation, high template library maintenance cost, and serious homogeneity of generation solutions. It is difficult to achieve deep integration of parametric modeling and image recognition, and cannot meet users' diverse customization needs.
A multi-scale feature pyramid network and hierarchical feature alignment strategy are adopted to achieve accurate conversion from pixel space to parameterized vector space through the joint mapping of multi-level visual semantic features of RGB images and physical size parameters, thus generating vector data of customized cabinets.
It achieves the rapid and accurate generation of customized cabinets, can make timely changes according to user needs, improves the flexibility and creativity of the design, meets various customization needs, and reduces calculation complexity.
Smart Images

Figure CN120747585A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of furniture technology, and more specifically, to a method, device, electronic device, and storage medium for generating a customized cabinet. Background Art
[0002] With the upgrading of personalized consumer demand, the home customization industry is gradually transforming towards digitalization. Since existing technologies generally adopt the traditional model of manual measurement, scheme design, rendering of renderings, and customer confirmation, customized homes face the following technical bottlenecks in the process of digital upgrading.
[0003] For example, there is a prominent contradiction between design efficiency and quality. Designers need to repeatedly draw two-dimensional sketches, build three-dimensional models and adjust dimensions before they can generate customized cabinet solutions. The whole process is time-consuming and labor-intensive, and manual drawing is prone to dimensional deviations. The cost of modifying the solution is high, and when customer needs change, modeling needs to start again from the basic parameters.
[0004] Although existing technologies attempt to integrate parametric modeling with image recognition, they have not yet achieved a deep integration of semantic parsing and engineering constraints, making it difficult for automated design systems to meet commercial standards.
[0005] Existing technology uses the network to identify the custom cabinet type in the reference image and then matches it with a preset template library. However, it can only identify explicit features and cannot extract specific home components. In addition, the maintenance cost of the template library is high and it is difficult to adapt to diverse design styles. The generated solutions are highly homogenized and lack creative design output. Summary of the Invention
[0006] The purpose of this application is to provide a method, device, electronic device and storage medium for generating customized cabinets, which can accurately and quickly generate vector data of customized cabinets, and timely change the customization plan according to user needs. It has higher flexibility and can achieve a deep integration of demand-based furniture customization and semantic analysis to meet users' various customization needs.
[0007] In a first aspect, an embodiment of the present application provides a method for generating a customized cabinet, the method comprising:
[0008] Acquire image data containing the cabinet structure;
[0009] Performing cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet;
[0010] Inputting the image data containing only the cabinet into the appearance prediction model for prediction to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories;
[0011] Inputting the image data containing only the cabinet into the interior space prediction model for prediction, and obtaining the interior space cabinet coordinates and the corresponding interior space prediction categories;
[0012] Performing size post-processing on the exterior cabinet coordinates and the interior cabinet coordinates to obtain standard exterior cabinet coordinates and standard interior cabinet coordinates;
[0013] Generate vector customized cabinet data based on standard appearance cabinet coordinates, appearance prediction category and standard interior space cabinet coordinates, interior space prediction category.
[0014] In the above implementation process, detection, prediction, size post-processing and other operations are performed based on the image data containing the cabinet body to generate vector data of the customized cabinet. The vector data of the customized cabinet can be generated accurately and quickly, and the customization plan can be changed in time according to user needs. It has higher flexibility and can achieve a deep integration of demand-based furniture customization and semantic analysis to meet users' various customization needs.
[0015] Furthermore, the step of inputting the image data containing only the cabinet into the appearance prediction model for prediction to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories includes:
[0016] Inputting the image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to perform feature extraction to obtain a first feature map;
[0017] Inputting the first feature map into an encoder of the appearance prediction model to obtain an enhanced first feature map;
[0018] Inputting the enhanced first feature map into a decoder of the appearance prediction model to obtain appearance polygon coordinates and appearance polygon categories;
[0019] Loss function training is performed on the appearance polygon coordinates and the appearance polygon categories to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories.
[0020] In the above implementation process, the image is processed in sequence through feature extraction, encoder, decoder, and loss training according to the appearance prediction model. The cabinets in the image can be separated one by one to obtain a complete cabinet with an appearance that meets user needs, and the corresponding category can be identified, which saves customization time, facilitates changes, and is more flexible and changeable.
[0021] Furthermore, the step of inputting the image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to perform feature extraction to obtain a first feature map includes:
[0022] Converting the image data containing only the cabinet into a grayscale image;
[0023] Extracting pixel features from the grayscale image;
[0024] Performing multi-scale mapping on the pixel features according to three convolutional layers to flatten the pixel features into a feature sequence;
[0025] A position code is added to each pixel feature in the feature sequence to obtain the first feature map.
[0026] In the above implementation process, the image data containing only the cabinet is converted into a grayscale image and then multi-scale mapping and position encoding are added. This can maintain the order of the feature sequence, improve the accuracy of the features, and facilitate subsequent predictions.
[0027] Furthermore, the step of adding a position code to each pixel feature in the feature sequence to obtain the first feature map includes:
[0028] Adding sine position coding and cosine position coding to each pixel feature in the feature sequence respectively;
[0029] The pixel features with added sine position coding and cosine position coding are mapped and connected to obtain the first feature map.
[0030] In the above implementation process, sine position coding and cosine position coding are added to the feature sequence and mapped and connected, which can clarify the position of pixel features in the feature sequence, realize indexing of different dimensions, and improve the model's recognition rate of features.
[0031] Furthermore, the step of inputting the enhanced first feature map into the decoder of the appearance prediction model to obtain the appearance polygon coordinates and the appearance polygon category includes:
[0032] The decoder of the appearance prediction model includes multiple decoding module layers, each decoding module layer includes a self-attention module, a multi-scale deformable cross attention module and a feedforward network module;
[0033] Each decoding module layer receives the enhanced first feature map;
[0034] The first decoding module layer generates a random appearance polygon sequence after receiving the enhanced first feature map;
[0035] The other decoding module layers except the first decoding module layer receive the appearance polygon sequence output by the previous decoding module layer through the self-attention module, and obtain the output of the corresponding self-attention module. The output of the self-attention module and the enhanced first feature Figure 1 The multi-scale deformable cross attention module is input, and then the appearance polygon coordinates and the appearance polygon category are obtained through the feedforward network module.
[0036] In the above implementation process, the first feature map is processed according to multiple decoding module layers in the decoder. Through interaction in the self-attention module, the multi-scale deformable cross-attention module focuses on different areas of the input first feature map, thereby achieving accurate positioning and prediction of the cabinet and improving the prediction accuracy.
[0037] Furthermore, the step of inputting the image data containing only the cabinet into the interior space prediction model for prediction to obtain the interior space cabinet coordinates and their corresponding interior space prediction categories includes:
[0038] Inputting the image data containing only the cabinet into the feature extraction backbone network of the interior space prediction model to perform feature extraction, thereby obtaining a second feature map;
[0039] Inputting the second feature map into the encoder of the inner space prediction model to obtain an enhanced second feature map;
[0040] Inputting the enhanced second feature map into the decoder of the inner space prediction model to obtain inner space polygon coordinates and inner space polygon categories;
[0041] Loss function training is performed on the inner space polygon coordinates and the inner space polygon categories to obtain the inner space cabinet coordinates and the corresponding inner space prediction categories.
[0042] In the above implementation process, the internal space structure of the cabinet is predicted through the internal space prediction model, which can accurately display the internal space structure, avoid errors caused by manual mapping, reduce the difficulty of measuring the internal space structure, and improve the efficiency of customized cabinet generation.
[0043] Furthermore, the step of generating vector customized cabinet data according to the standard appearance cabinet coordinates, appearance prediction category and the standard interior space cabinet coordinates, interior space prediction category includes:
[0044] Generating topological relationship data including cabinet layout according to the standard appearance cabinet coordinates, appearance prediction category and standard interior space cabinet coordinates, interior space prediction category;
[0045] The vector customized cabinet data is generated according to the topological relationship data including the cabinet layout.
[0046] In the above implementation process, the topological relationship between the cabinet appearance and the interior space is generated according to the standard appearance cabinet coordinates, appearance prediction category, standard interior space cabinet coordinates, and interior space prediction category. The cabinet can be fully displayed, and the cabinet structure and size can be clearly presented, which facilitates the modification and improvement of the cabinet.
[0047] In a second aspect, an embodiment of the present application further provides a device for generating a customized cabinet, the device comprising:
[0048] An acquisition module, used for acquiring image data containing the cabinet structure;
[0049] A cabinet detection module, configured to perform cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet;
[0050] A prediction module, configured to input the image data containing only the cabinet into an appearance prediction model for prediction, thereby obtaining the appearance cabinet coordinates and their corresponding appearance prediction categories; and further configured to input the image data containing only the cabinet into an interior space prediction model for prediction, thereby obtaining the interior space cabinet coordinates and their corresponding interior space prediction categories;
[0051] A size post-processing module, configured to perform size post-processing on the exterior cabinet coordinates and the interior cabinet coordinates to obtain standard exterior cabinet coordinates and standard interior cabinet coordinates;
[0052] The generation module is used to generate vector customized cabinet data based on standard appearance cabinet coordinates, appearance prediction categories and standard interior cabinet coordinates, interior prediction categories.
[0053] In the above implementation process, detection, prediction, size post-processing and other operations are performed based on the image data containing the cabinet body to generate vector data of the customized cabinet. The vector data of the customized cabinet can be generated accurately and quickly, and the customization plan can be changed in time according to user needs. It has higher flexibility and can achieve a deep integration of demand-based furniture customization and semantic analysis to meet users' various customization needs.
[0054] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in any one of the first aspects when executing the computer program.
[0055] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which instructions are stored. When the instructions are executed on a computer, the computer executes the method as described in any one of the first aspects.
[0056] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by practicing the above-mentioned technology of the present disclosure.
[0057] It can be implemented according to the contents of the specification. The following is a detailed description of the preferred embodiments of the present application with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the range values. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0059] Figure 1 A flowchart of a method for generating a customized cabinet is provided for an embodiment of the present application;
[0060] Figure 2 A schematic diagram of the structure of a device for generating a customized cabinet is provided for an embodiment of the present application;
[0061] Figure 3 A schematic diagram of the structural composition of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0063] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0064] The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0065] Existing technologies are unable to integrate parametric modeling with image recognition, or to achieve a deep integration of semantic analysis and engineering constraints. This results in automated design systems lacking fixed usage standards and unable to provide standardized customization for individual user needs. Cabinet customization based on template libraries also fails to extract specific home components, resulting in a rigid style and a high degree of homogeneity in the resulting designs, lacking creative output.
[0066] In response to the defects of the above-mentioned existing technologies, this application proposes a hierarchical feature alignment strategy based on a multi-scale feature pyramid network. By jointly mapping the multi-level visual semantic features of RGB images with physical size parameters, it achieves accurate conversion from pixel space to parameterized space to vector space, reduces computational complexity, and realizes standardized and various customization requirements.
[0067] Example 1
[0068] Figure 1This is a flow chart of a method for generating a customized cabinet provided in an embodiment of the present application. Figure 1 As shown, the method includes:
[0069] S1, acquiring image data containing the cabinet structure;
[0070] S2, performing cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet;
[0071] S3, inputting the image data containing only the cabinet into the appearance prediction model for prediction, and obtaining the appearance cabinet coordinates and their corresponding appearance prediction categories;
[0072] S4, inputting the image data containing only the cabinet into the interior space prediction model for prediction, and obtaining the interior space cabinet coordinates and their corresponding interior space prediction categories;
[0073] S5, performing dimensional post-processing on the exterior cabinet coordinates and the interior cabinet coordinates to obtain standard exterior cabinet coordinates and standard interior cabinet coordinates;
[0074] S6, generating vector customized cabinet data according to the standard appearance cabinet coordinates, appearance prediction category and the standard interior space cabinet coordinates, interior space prediction category.
[0075] In the above implementation process, detection, prediction, size post-processing and other operations are performed based on the image data containing the cabinet body to generate vector data of the customized cabinet. The vector data of the customized cabinet can be generated accurately and quickly, and the customization plan can be changed in time according to user needs. It has higher flexibility and can achieve a deep integration of demand-based furniture customization and semantic analysis to meet users' various customization needs.
[0076] In S1, the image data containing the cabinet structure in this application is RGB image data.
[0077] In S2, since the RGB image data may contain complex backgrounds, and these background contents will cause great interference to the actual layout recognition of the cabinet, this application needs to use a pre-built detection model to separate the cabinet part from the complex background, and then apply it to the actual cabinet customization process. The detection model uses the yolov8 open source model, the input is RGB image data containing the cabinet structure, and the output is image data without background and only containing the cabinet.
[0078] Furthermore, S3 includes:
[0079] Inputting image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to extract features and obtain a first feature map;
[0080] Inputting the first feature map into an encoder of an appearance prediction model to obtain an enhanced first feature map;
[0081] Input the enhanced first feature map into the decoder of the appearance prediction model to obtain the appearance polygon coordinates and appearance polygon category;
[0082] The loss function is trained on the appearance polygon coordinates and appearance polygon categories to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories.
[0083] In the above implementation process, the image is processed in sequence through feature extraction, encoder, decoder, and loss training according to the appearance prediction model. The cabinets in the image can be separated one by one to obtain a complete cabinet with an appearance that meets user needs, and the corresponding category can be identified, which saves customization time, facilitates changes, and is more flexible and changeable.
[0084] In this application, the feature extraction backbone network adopts CNNBackbone, the encoder is a Transformer encoder, and the decoder is a Transformer decoder. The encoder of the embodiment of this application takes the first feature map as input and outputs the enhanced first feature map of the same resolution. The encoder consists of multiple encoding layers, and each encoding layer consists of a multi-scale deformable cross attention module and a feedforward network module.
[0085] To avoid the computational and memory complexity of the standard Transformer, this application adopts a multi-scale deformable cross attention module so that the feature map only focuses on the surrounding key points instead of all spatial positions on the feature map.
[0086] The training loss functions include: vertex label classification loss and vertex coordinate regression loss. The vertex coordinate regression loss is used for loss training of the appearance polygon coordinates, and the vertex label classification loss is used for loss training of the appearance polygon category. The vertex coordinate regression loss in this application adopts the mean absolute error loss function, and the vertex label classification loss adopts the cross quotient loss function.
[0087] Furthermore, the step of inputting the image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to extract features and obtain a first feature map includes:
[0088] Convert the image data containing only the cabinet into a grayscale image;
[0089] Extract pixel features from grayscale images;
[0090] Multi-scale mapping is performed on pixel features according to three convolutional layers to flatten the pixel features into feature sequences;
[0091] A position code is added to each pixel feature in the feature sequence to obtain the first feature map.
[0092] In the above implementation process, the image data containing only the cabinet is converted into a grayscale image and then multi-scale mapping and position encoding are added. This can maintain the order of the feature sequence, improve the accuracy of the features, and facilitate subsequent predictions.
[0093] The image data containing only the cabinet is converted into a grayscale image through OpenCV. Since accurate coordinate point positioning and obtaining the alignment order require layout and global features, three convolution layers are used for multi-scale mapping in the embodiment of the present application, and each pixel feature is flattened into a feature sequence.
[0094] Furthermore, the step of adding a position code to each pixel feature in the feature sequence to obtain a first feature map includes:
[0095] Add sine position coding and cosine position coding to each pixel feature in the feature sequence respectively;
[0096] The pixel features with added sine position coding and cosine position coding are mapped and connected to obtain the first feature map.
[0097] In the above implementation process, sine position coding and cosine position coding are added to the feature sequence and mapped and connected, which can clarify the position of pixel features in the feature sequence, realize indexing of different dimensions, and improve the model's recognition rate of features.
[0098] In order to maintain the order of the input feature sequence, sine position encoding / cosine position encoding is added to each pixel feature position (corresponding to the addition of feature vectors), and all feature maps are concatenated as the input of the Transformer encoder. The formula is as follows:
[0099]
[0100] Among them, PE is the position code, d model is the dimension of the feature vector corresponding to the pixel feature, pos is the position of the current pixel feature in the entire input pixel feature, i represents the vector dimension index of the position encoding, sin is the sine trigonometric function, and cos is the cosine trigonometric function.
[0101] Furthermore, the step of inputting the enhanced first feature map into a decoder of the appearance prediction model to obtain the appearance polygon coordinates and the appearance polygon category includes:
[0102] The decoder of the appearance prediction model contains multiple decoding module layers, each of which contains a self-attention module, a multi-scale deformable cross-attention module, and a feedforward network module;
[0103] Each decoding module layer receives the enhanced first feature map;
[0104] The first decoding module layer receives the enhanced first feature map and generates a random appearance polygon sequence;
[0105] The other decoding module layers except the first decoding module layer receive the appearance polygon sequence output by the previous decoding module layer through the self-attention module, and obtain the output of the corresponding self-attention module. The output of the self-attention module and the enhanced first feature Figure 1 It inputs the multi-scale deformable cross attention module and then passes through the feedforward network module to obtain the appearance polygon coordinates and the appearance polygon category.
[0106] In the above implementation process, the first feature map is processed according to multiple decoding module layers in the decoder. Through interaction in the self-attention module, the multi-scale deformable cross-attention module focuses on different areas of the input first feature map, thereby achieving accurate positioning and prediction of the cabinet and improving the prediction accuracy.
[0107] The decoder contains multiple decoding module layers, each layer contains a self-attention module, a multi-scale deformable cross-attention module and a feedforward network module. Each decoder receives an enhanced first feature map from the encoder and a set of M*N polygon sequences (appearance polygon sequences) from the previous layer. The polygon sequence is generated by the output of the feedforward network module. The polygon sequence of the first layer is randomly generated. M represents the maximum number of polygons in a cabinet, and N represents the number of vertices of each polygon. The polygon sequence first interacts with each other in the self-attention module. The multi-scale deformable cross-attention module focuses on different areas of the input first feature map. Finally, the output of the decoder is passed to the feedforward network module to predict the type and coordinates of each polygon. Specifically, the output is a vector (M, N, O), where M represents the maximum number of polygons in a cabinet, N represents the number of vertices of each polygon, and O represents the category and coordinates of each vertex. For example, in (c, x, y), c is the category and (x, y) is the vertex coordinate.
[0108] When predicting appearance, the appearance polygon categories are divided into 6 categories, namely, cell-enclosed space units, door panels, drawers, open space units, and glass doors, and each polygon is a rectangle.
[0109] Furthermore, S4 includes:
[0110] Input the image data containing only the cabinet into the feature extraction backbone network of the interior space prediction model to extract features and obtain a second feature map;
[0111] Inputting the second feature map into the encoder of the inner space prediction model to obtain an enhanced second feature map;
[0112] Input the enhanced second feature map into the decoder of the inner space prediction model to obtain the inner space polygon coordinates and inner space polygon categories;
[0113] The loss function is trained on the inner space polygon coordinates and inner space polygon categories to obtain the inner space cabinet coordinates and their corresponding inner space prediction categories.
[0114] In the above implementation process, the internal space structure of the cabinet is predicted through the internal space prediction model, which can accurately display the internal space structure, avoid errors caused by manual mapping, reduce the difficulty of measuring the internal space structure, and improve the efficiency of customized cabinet generation.
[0115] In an embodiment of the present application, the prediction of the internal space structure of the cabinet is the same as the prediction of the appearance, and the same model is used for prediction, that is, the structure of the appearance prediction model and the internal space prediction model is the same, but the training data in the model training process is different. The internal space prediction category is divided into 3 categories, namely, closed space units, drawers, and open space units, and each polygon is a rectangle.
[0116] In S5, the coordinates of the model prediction results may not be completely aligned. For example, the polygons of two adjacent door panels are not aligned up and down. Dimension post-processing is required based on the actual situation to convert the predicted exterior cabinet coordinates and interior cabinet coordinates into standard exterior cabinet coordinates and standard interior cabinet coordinates.
[0117] The difference between the two coordinates is determined to be within a preset threshold. If it is within the threshold, the maximum of the two values is manually set. The coordinate range of the input image data has been processed to be within the range of (0-255) during training data to ensure a more stable training process and faster convergence.
[0118] After receiving the corresponding predicted coordinates, the user needs to scale them proportionally according to the required size. The user can define the actual space size parameters according to the requirements.
[0119] Furthermore, S6 includes:
[0120] Generate topological relationship data including cabinet layout according to standard appearance cabinet coordinates, appearance prediction category and standard interior space cabinet coordinates, interior space prediction category;
[0121] Generate vector customized cabinet data based on topological relationship data including cabinet layout.
[0122] In the above implementation process, the topological relationship between the cabinet appearance and the interior space is generated according to the standard appearance cabinet coordinates, appearance prediction category, standard interior space cabinet coordinates, and interior space prediction category. The cabinet can be fully displayed, and the cabinet structure and size can be clearly presented, which facilitates the modification and improvement of the cabinet.
[0123] Based on the above-mentioned output of standard appearance cabinet coordinates, appearance prediction categories and standard interior cabinet coordinates, interior prediction categories, topological relationship data containing the cabinet layout is generated. The topological relationship is calculated based on the polygons predicted by the appearance and the polygons predicted by the interior. For example, a door panel polygon predicted by the appearance may contain multiple interior polygons. This inclusion relationship calculation can constitute the topological relationship of the cabinet layout, material properties and process parameters, and docking design software to identify data.
[0124] The embodiment of the present application is based on a multi-scale feature pyramid network and combines sine-cosine position encoding to achieve geometric embedding of image spatial position information, solving the problem of feature loss caused by scale changes in the existing technology; a hierarchical feature alignment strategy is proposed, which realizes the precise conversion of pixel space to parameterized vector space through the joint mapping of multi-level visual semantic features (color, texture, structure) of RGB images and physical size parameters; the computational complexity is reduced through multi-scale deformable attention, and result sequence prediction is introduced to directly output polygon vertex coordinates and category labels through iterative optimization.
[0125] Example 2
[0126] In order to execute the method corresponding to the above embodiment 1 to achieve the corresponding functions and technical effects, a device for generating a customized cabinet is provided below, such as Figure 2 As shown, the device includes:
[0127] Acquisition module 1, used to acquire image data containing cabinet structure;
[0128] The cabinet detection module 2 is used to perform cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet;
[0129] Prediction module 3 is used to input image data containing only the cabinet into the appearance prediction model for prediction, and obtain the appearance cabinet coordinates and their corresponding appearance prediction categories; it is also used to input image data containing only the cabinet into the interior space prediction model for prediction, and obtain the interior space cabinet coordinates and their corresponding interior space prediction categories;
[0130] The size post-processing module 4 is used to perform size post-processing on the external cabinet coordinates and the internal cabinet coordinates to obtain the standard external cabinet coordinates and the standard internal cabinet coordinates;
[0131] The generating module 5 is used to generate vector customized cabinet data according to the standard appearance cabinet coordinates, appearance prediction category and the standard interior space cabinet coordinates, interior space prediction category.
[0132] In the above implementation process, detection, prediction, size post-processing and other operations are performed based on the image data containing the cabinet body to generate vector data of the customized cabinet. The vector data of the customized cabinet can be generated accurately and quickly, and the customization plan can be changed in time according to user needs. It has higher flexibility and can achieve a deep integration of demand-based furniture customization and semantic analysis to meet users' various customization needs.
[0133] Furthermore, the prediction module 3 is also used for:
[0134] Inputting image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to extract features and obtain a first feature map;
[0135] Inputting the first feature map into an encoder of an appearance prediction model to obtain an enhanced first feature map;
[0136] Input the enhanced first feature map into the decoder of the appearance prediction model to obtain the appearance polygon coordinates and appearance polygon category;
[0137] The loss function is trained on the appearance polygon coordinates and appearance polygon categories to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories.
[0138] In the above implementation process, the image is processed in sequence through feature extraction, encoder, decoder, and loss training according to the appearance prediction model. The cabinets in the image can be separated one by one to obtain a complete cabinet with an appearance that meets user needs, and the corresponding category can be identified, which saves customization time, facilitates changes, and is more flexible and changeable.
[0139] Furthermore, the prediction module 3 is also used for:
[0140] Convert the image data containing only the cabinet into a grayscale image;
[0141] Extract pixel features from grayscale images;
[0142] Multi-scale mapping is performed on pixel features according to three convolutional layers to flatten the pixel features into feature sequences;
[0143] A position code is added to each pixel feature in the feature sequence to obtain the first feature map.
[0144] In the above implementation process, the image data containing only the cabinet is converted into a grayscale image and then multi-scale mapping and position encoding are added. This can maintain the order of the feature sequence, improve the accuracy of the features, and facilitate subsequent predictions.
[0145] Furthermore, the prediction module 3 is also used for:
[0146] Add sine position coding and cosine position coding to each pixel feature in the feature sequence respectively;
[0147] The pixel features with added sine position coding and cosine position coding are mapped and connected to obtain the first feature map.
[0148] In the above implementation process, sine position coding and cosine position coding are added to the feature sequence and mapped and connected, which can clarify the position of pixel features in the feature sequence, realize indexing of different dimensions, and improve the model's recognition rate of features.
[0149] Furthermore, the prediction module 3 is also used for:
[0150] The decoder of the appearance prediction model contains multiple decoding module layers, each of which contains a self-attention module, a multi-scale deformable cross-attention module, and a feedforward network module;
[0151] Each decoding module layer receives the enhanced first feature map;
[0152] The first decoding module layer receives the enhanced first feature map and generates a random appearance polygon sequence;
[0153] The other decoding module layers except the first decoding module layer receive the appearance polygon sequence output by the previous decoding module layer through the self-attention module, and obtain the output of the corresponding self-attention module. The output of the self-attention module and the enhanced first feature Figure 1 It inputs the multi-scale deformable cross attention module and then passes through the feedforward network module to obtain the appearance polygon coordinates and the appearance polygon category.
[0154] In the above implementation process, the first feature map is processed according to multiple decoding module layers in the decoder. Through interaction in the self-attention module, the multi-scale deformable cross-attention module focuses on different areas of the input first feature map, thereby achieving accurate positioning and prediction of the cabinet and improving the prediction accuracy.
[0155] Furthermore, the prediction module 3 is also used for:
[0156] Input the image data containing only the cabinet into the feature extraction backbone network of the interior space prediction model to extract features and obtain a second feature map;
[0157] Inputting the second feature map into the encoder of the inner space prediction model to obtain an enhanced second feature map;
[0158] Input the enhanced second feature map into the decoder of the inner space prediction model to obtain the inner space polygon coordinates and inner space polygon categories;
[0159] The loss function is trained on the inner space polygon coordinates and inner space polygon categories to obtain the inner space cabinet coordinates and their corresponding inner space prediction categories.
[0160] In the above implementation process, the internal space structure of the cabinet is predicted through the internal space prediction model, which can accurately display the internal space structure, avoid errors caused by manual mapping, reduce the difficulty of measuring the internal space structure, and improve the efficiency of customized cabinet generation.
[0161] Furthermore, the generating module 5 is further configured to:
[0162] Generate topological relationship data including cabinet layout according to standard appearance cabinet coordinates, appearance prediction category and standard interior space cabinet coordinates, interior space prediction category;
[0163] Generate vector customized cabinet data based on topological relationship data including cabinet layout.
[0164] In the above implementation process, the topological relationship between the cabinet appearance and the interior space is generated according to the standard appearance cabinet coordinates, appearance prediction category, standard interior space cabinet coordinates, and interior space prediction category. The cabinet can be fully displayed, and the cabinet structure and size can be clearly presented, which facilitates the modification and improvement of the cabinet.
[0165] The above-mentioned device for generating a customized cabinet can implement the method of the above-mentioned embodiment 1. The options in the above-mentioned embodiment 1 are also applicable to this embodiment and will not be described in detail here.
[0166] The rest of the contents of the embodiments of this application can refer to the contents of the above-mentioned embodiment 1, and will not be repeated in this embodiment.
[0167] Example 3
[0168] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the method for generating a customized cabinet in embodiment 1.
[0169] Optionally, the above-mentioned electronic device may be a server.
[0170] See Figure 3 , Figure 3 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may include a processor 31, a communication interface 32, a memory 33, and at least one communication bus 34. The communication bus 34 is used to enable direct communication between these components.
[0171] Optionally, the electronic device may further include a storage controller and an input / output unit. The memory 33, storage controller, processor 31, peripheral interface, and input / output unit are electrically connected to each other directly or indirectly to achieve data transmission or interaction.
[0172] The input and output unit is used to provide users with the ability to create tasks and to create optional start time periods or preset execution times for the tasks to enable interaction between the user and the server. The input and output unit can be, but is not limited to, a mouse and keyboard.
[0173] I understand. Figure 3 The structure shown is only for illustration, and the electronic device may also include Figure 3 More or fewer components than shown, or with Figure 3 In addition, the embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for generating a custom cabinet according to the first embodiment.
[0174] An embodiment of the present application further provides a computer program product, which, when running on a computer, enables the computer to execute the method described in the method embodiment.
[0175] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included within the scope of protection of the present application. It should be noted that similar numbers and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.
[0176] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included within the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for generating a customized cabinet, characterized in that: The method comprises: Acquire image data containing the cabinet structure; Performing cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet; Inputting the image data containing only the cabinet into the appearance prediction model for prediction to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories; Inputting the image data containing only the cabinet into the interior space prediction model for prediction, and obtaining the interior space cabinet coordinates and the corresponding interior space prediction categories; Performing size post-processing on the exterior cabinet coordinates and the interior cabinet coordinates to obtain standard exterior cabinet coordinates and standard interior cabinet coordinates; Vector customized cabinet data is generated according to the standard appearance cabinet coordinates, the appearance prediction category and the standard interior space cabinet coordinates, and the interior space prediction category.
2. The method for generating a customized cabinet according to claim 1, characterized in that: The step of inputting the image data containing only the cabinet into the appearance prediction model for prediction to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories includes: Inputting the image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to perform feature extraction to obtain a first feature map; Inputting the first feature map into an encoder of the appearance prediction model to obtain an enhanced first feature map; Inputting the enhanced first feature map into a decoder of the appearance prediction model to obtain appearance polygon coordinates and appearance polygon categories; Loss function training is performed on the appearance polygon coordinates and the appearance polygon categories to obtain the appearance cabinet coordinates and their corresponding appearance prediction categories.
3. The method for generating a customized cabinet according to claim 2, characterized in that: The step of inputting the image data containing only the cabinet into the feature extraction backbone network of the appearance prediction model to perform feature extraction to obtain a first feature map includes: Converting the image data containing only the cabinet into a grayscale image; Extracting pixel features from the grayscale image; Performing multi-scale mapping on the pixel features according to three convolutional layers to flatten the pixel features into a feature sequence; A position code is added to each pixel feature in the feature sequence to obtain the first feature map.
4. The method for generating a customized cabinet according to claim 3, characterized in that: The step of adding a position code to each pixel feature in the feature sequence to obtain the first feature map includes: Adding sine position coding and cosine position coding to each pixel feature in the feature sequence respectively; The pixel features with added sine position coding and cosine position coding are mapped and connected to obtain the first feature map.
5. The method for generating a customized cabinet according to claim 2, characterized in that: The step of inputting the enhanced first feature map into the decoder of the appearance prediction model to obtain the appearance polygon coordinates and the appearance polygon category includes: The decoder of the appearance prediction model includes multiple decoding module layers, each decoding module layer includes a self-attention module, a multi-scale deformable cross attention module and a feedforward network module; Each decoding module layer receives the enhanced first feature map; The first decoding module layer generates a random appearance polygon sequence after receiving the enhanced first feature map; The other decoding module layers except the first decoding module layer receive the appearance polygon sequence output by the previous decoding module layer through the self-attention module to obtain the output of the corresponding self-attention module. The output of the self-attention module and the enhanced first feature map are input into the multi-scale deformable cross attention module together, and then pass through the feedforward network module to obtain the appearance polygon coordinates and the appearance polygon category.
6. The method for generating a customized cabinet according to claim 1, characterized in that: The step of inputting the image data containing only the cabinet into the interior space prediction model for prediction to obtain the interior space cabinet coordinates and the corresponding interior space prediction category includes: Inputting the image data containing only the cabinet into the feature extraction backbone network of the interior space prediction model to perform feature extraction, thereby obtaining a second feature map; Inputting the second feature map into the encoder of the inner space prediction model to obtain an enhanced second feature map; Inputting the enhanced second feature map into the decoder of the inner space prediction model to obtain inner space polygon coordinates and inner space polygon categories; Loss function training is performed on the inner space polygon coordinates and the inner space polygon categories to obtain the inner space cabinet coordinates and the corresponding inner space prediction categories.
7. The method for generating a customized cabinet according to claim 1, characterized in that: The step of generating vector customized cabinet data according to the standard appearance cabinet coordinates, the appearance prediction category and the standard interior space cabinet coordinates, and the interior space prediction category includes: Generating topological relationship data including cabinet layout according to the standard appearance cabinet coordinates, the appearance prediction category and the standard interior space cabinet coordinates, and the interior space prediction category; The vector customized cabinet data is generated according to the topological relationship data including the cabinet layout.
8. A device for generating a customized cabinet, characterized in that: The device comprises: An acquisition module, used for acquiring image data containing the cabinet structure; A cabinet detection module, configured to perform cabinet detection on the image data containing the cabinet structure according to a pre-built detection model to obtain image data containing only the cabinet; A prediction module, configured to input the image data containing only the cabinet into an appearance prediction model for prediction, thereby obtaining the appearance cabinet coordinates and their corresponding appearance prediction categories; and further configured to input the image data containing only the cabinet into an interior space prediction model for prediction, thereby obtaining the interior space cabinet coordinates and their corresponding interior space prediction categories; A size post-processing module, configured to perform size post-processing on the exterior cabinet coordinates and the interior cabinet coordinates to obtain standard exterior cabinet coordinates and standard interior cabinet coordinates; The generation module is used to generate vector customized cabinet data based on standard appearance cabinet coordinates, appearance prediction categories and standard interior cabinet coordinates, interior prediction categories.
9. An electronic device, characterized in that: It includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the method for generating a customized cabinet according to any one of claims 1 to 7.
10. A storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the method for generating a customized cabinet as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Contour extraction method and device, electronic equipment and storage medium
CN112183541A
Cabinet body structure identification method and device based on single image, equipment and medium
CN113744350A
Decorative drawing generation method and device of customized cabinet
CN116136919A
Cabinet image correction method and device, computer equipment, readable storage medium and program product
CN119206208A
Cabinet model generation method and system
CN119359957A