METHOD FOR IDENTIFYING A PART USING A CONVOLUTIONAL NEURAL NETWORK
By employing a modified convolutional neural network architecture and procedural generation techniques, the method effectively identifies different types of rooms with a high recognition rate, overcoming the challenges faced by existing neural network algorithms in distinguishing between similar rooms.
Patent Information
- Application Number
- FR2023004786
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-15
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-05-15
AI Technical Summary
Existing neural network algorithms struggle to accurately distinguish between nearby rooms, such as kitchens and bathrooms, due to similarities in furniture dimensions and water points, resulting in a high error rate.
A convolutional neural network architecture is developed, inspired by the 'VGG 19' configuration but with two convolution layers removed, along with the use of procedural generation to enhance training data, particularly for identifying rooms like bedrooms, kitchens, bathrooms, and living rooms.
The proposed method achieves a high recognition rate for identifying different types of rooms, significantly improving upon the performance of 'VGG 16' and 'VGG 19' configurations, and addresses the limitations of existing technologies by incorporating procedural generation to enrich training data.
Smart Images

Figure 00000020_0000 
Figure 00000020_0001 
Figure 00000021_0000
Abstract
Description
Title of the invention: METHOD FOR IDENTIFYING A PART BY MEANS OF A CONVOLUTIONAL NEURAL NETWORK Field of invention
[0001] The invention relates to a method for identifying a room, for example a room in an apartment, a house or an industrial premises. To do this, the invention relates more specifically to a method for identifying a room using a convolutional neural network, i.e. a specific artificial intelligence algorithm for identifying an image or a stream of images.
[0002] The aim of the invention is to limit the error rate in identifying a type of part.
[0003] The invention can be applied in all fields for which it is desired to identify a room, for example to guide installers, technicians or robots. Thus, the invention can be used to guide the movements of a robot between several rooms of a factory or even to detect the room in which a technician is positioned.
[0004] The invention finds a particularly advantageous application for automatically detecting the type of rooms in a home and facilitating the assessment of claims for an expert. State of the art
[0005] To automatically identify a part using an image or a stream of images from a camera, for example captured by a smartphone, it is possible to use neural networks.
[0006] Indeed, neural networks are now known for their identification capacity. By distinguishing characteristic objects of a room, these neural networks should be effective in characterizing a room and, thus, recognizing whether an expert is positioned in one room or another of a home.
[0007] For example, by detecting a bed, a chest of drawers, two bedside tables, the neural networks should be able to detect a bedroom compared to a living room in which a coffee table and a sofa would be detected.
[0008] However, the neural network algorithms tested in the context of the invention prove to be ineffective in distinguishing between nearby rooms. For example, the algorithms tested have a high error rate when it comes to differentiating a kitchen from a bathroom. Indeed, these two rooms include furniture with similar dimensions as well as at least one water point, such as a sink.
[0009] It is therefore sought to obtain a network of particular neurons making it possible to improve the detection of rooms, in particular the detection of rooms in a house or an apartment.
[0010] An identification neural network is a mathematical structure in which an input image is compared to training images to detect whether or not the input image matches one of the labels associated with the training images.
[0011] For example, if the aim is to differentiate a kitchen from a bathroom, images corresponding to bathrooms and images corresponding to kitchens are used to train a neural network. Then, when an image is placed as input to the neural network, the latter must be able to distinguish whether the presented image corresponds to a bathroom, a kitchen or does not correspond to any of the trained images.
[0012] There is a very wide variety of neural networks but neural networks are generally constructed with a set of neurons interconnected by layers, in which the neurons of each layer perform the same mathematical function. In addition, each neuron is typically associated with an activation function in which one or more parameters can be modified during the learning phase.
[0013] This activation function can take a large number of forms and, in the literature, we find activation functions of the “SIGMOID” type, of the “TANH” type, of the “RELU” type, of the “LEAKY RELU” type, of the “MAX OUT” type, or even of the “ELU” type, to name but a few.
[0014] Thus, to build a neural network, it is necessary to define the mathematical functions of each neuron, the activation functions, as well as the interconnections between its neurons. Then, it is possible to carry out the learning phase and verify that the expected results of this neural network meet the needs of the application. Given the increasing number of neurons used in neural networks, the multiple possibilities of interconnection and the multiple possibilities of defining the mathematical functions and the activation functions, it is not possible to test all the possibilities of building a neural network for a specific application.
[0015] To simplify this work of constructing a neural network, several families of neural networks have been developed during numerous research projects and it appears that convolutional neural networks are particularly effective for identification needs, for example to distinguish a cat from a dog or to recognize handwritten characters.
[0016] These convolution neural networks integrate layers which carry out mathematical convolution functions on the images present at the input of these layers. For each layer of neurons performing a convolution, the size of the convolution filter applied to the input image must be defined so that a first neuron performs a convolution on a first image of the input image, a second neuron performs a convolution on a second image of the input image, and so on to cover the entire input image.
[0017] In addition to the size of the convolution filter, it is also necessary to define how to manage the movements between two filters, also called "stride" in English literature, and the resizing of the images resulting from these filters, also called "padding" in English literature.
[0018] The displacement between two filters can be defined by the number of pixel shifts between the application of one filter and the application of the next filter.
[0019] Resizing can be used or not. When it is not used, the convolution layer is defined by a "VALID padding" and the images coming from the convolution layer are of sizes smaller than those present at the input of this layer.
[0020] When resizing is used, the convolution layer is defined by a "S AME padding" and the resizing aims to give as much influence to the border pixels as to the other pixels. This resizing helps to limit the edge effect. Typically, a pixel located on the edge of an image will be used for fewer neurons in the convolution layer than the other pixels. The edge pixels will therefore have less influence on the set of output images. To limit this edge effect, the resizing increases the size of the initial matrix with zeros all around the edge pixels, so that these pixels are not underrepresented.
[0021] Thus, for a convolution layer, it is possible to define the set of neurons by selecting: - the size of the filter, - moving between two filters, - resizing, and - the activation function.
[0022] This convolution layer typically allows for detecting salient features in the input image of the convolution layer. However, using a single convolution layer is insufficient for characterizing complex objects, such as different types of parts.
[0023] To do this, it is necessary to use multiple convolution layers placed one after the other to detect salient features on previously detected salient features, and so on. However, using a convolution layer generates an increase in the volume of data because each filter of the convolution layer generates an image at the output of the convolution layer. convolution. By using a second convolution layer after a first convolution layer, the second convolution layer has neurons to process all the thumbnails from the first convolution layer, so the multiplication of convolution layers leads to a significant increase in data volume and processing time.
[0024] To reduce this volume of data between two convolution layers, it is known to use pooling layers, also called “pooling” in English literature.
[0025] These grouping layers make it possible to limit the volume of data by applying a filter to the output images of the convolution layers and by keeping a single pixel for all the pixels present in the filter. Once again, when it is chosen to use a grouping layer, several grouping possibilities are possible.
[0026] Indeed, it is necessary to select: - the size of the filter, - the mathematical grouping function, for example by detecting the average value of the pixels or the maximum value, - moving between two filters, and - resizing.
[0027] In addition to these convolution and pooling layers, there are also a large number of potentially usable layers, such as interconnection layers also called "fully connected" in the English literature. These interconnection layers make it possible to interconnect all the results previously obtained by the neurons of the convolution or pooling layers.
[0028] To build a convolution neural network, it is therefore necessary to define the size of the input images, the number of convolution, grouping and interconnection layers and to configure them with the parameters previously described.
[0029] Given the multiple possible implementations for all these choices, it is not possible to test all the possibilities for constructing a convolutional neural network for a specific application.
[0030] To facilitate the implementation of a convolutional neural network, a large number of standards have been defined in scientific publications so that a person skilled in the art can detect which neural network would be a priori relevant for his application.
[0031] In the specific case of object identification, VGG type neural networks, for “Visual Geometry Group” in English literature, present particularly interesting performances.
[0032] These VGG-type neural networks integrate fixed characteristics, such as the size of the input images in 224 by 224 pixels and predefined parameters for the convolution and pooling layers. More precisely, as illustrated in the table in [Fig.l], six standardized structures were defined depending on the desired implementation complexity.
[0033] The first column illustrates configuration A in which: - a first convolution layer conv3-64 is followed by a maxpool pooling layer; - itself followed by a second convolution layer conv3-128; - also followed by a maxpool pooling layer; - then two successive convolution layers conv3-256; - followed by a maxpool pooling layer; - as well as two other successive convolution layers conv3-512; - followed by a maxpool pooling layer and two final successive convolution layers conv3-512; - followed by a maxpool grouping layer; three interconnection layers FC-4096, FC-4096, FC-1000 and a final Soft-max identification layer.
[0034] In this configuration, all the convolution layers conv3-64, conv3-128, conv3-256, conv3-512 are made with identical filters of 3 pixels by 3 pixels, a displacement between two filters of one pixel, a resizing of type "SAME padding" and using activation functions of type "RELU". In addition, the maxpool pooling layers use filters of 2 pixels by 2 pixels, a function for calculating the pooling pixel using the maximum, a displacement between two filters of two pixels and a resizing of type "VALID padding".
[0035] Thus, this particular structure makes it possible to reuse identical neurons for the different convolution layers conv3-64, conv3-128, conv3-256, conv3-512 and maxpool pooling even if the size of the output images of the convolution layers differs depending on the size of the input images, i.e. the number of convolution and pooling layers previously implemented.
[0036] A variant of this first configuration A, called A-LRN, is illustrated in the second column and proposes to use a local normalization layer LNR after the first convolution layer conv3-64.
[0037] Configuration B illustrated in the third column has thirteen processing layers, i.e. thirteen convolution layers conv3-64, conv3-128, conv3-256, conv3-512 and interconnection layers FC-4096, FC-4096 and FC-1000, unlike configurations A and A-LRN which only integrate eleven processing layers.
[0038] Thus, for configuration B, the input image is transmitted to two layers of successive convolution conv3-64 and two successive convolution conv3-128 layers are also placed after the first maxpooling layer.
[0039] Configuration C integrates sixteen processing layers by adding a specific convolution layer convl-256, convl-512 with a one-pixel-by-one-pixel filter before the last three maxpool pooling layers. Configuration C also uses three convolution layers conv3-256, conv3-512, convl-256, convl-512 before the last three maxpool pooling layers. The other convolution layers conv3-64, conv3-128 correspond to the convolution layers previously used in configuration B.
[0040] Configuration D is particularly well known since it corresponds to the widely used “VGG 16” configuration. This configuration D integrates sixteen processing layers with: - two conv3-64 convolution layers; - followed by a maxpool pooling layer; - followed by two convolution layers conv3-128; - followed by a maxpool pooling layer; - followed by three convolution layers conv3-256; - followed by a maxpool pooling layer; - followed by three convolution layers conv3-512; - followed by a maxpool pooling layer; - followed by three convolution layers conv3-512; - followed by a maxpool pooling layer; - followed by three interconnection layers FC-4096, FC-4096 and FC-1000 and a Soft-max identification layer.
[0041] Configuration E is also known since it corresponds to the “VGG 19” configuration. This configuration E integrates nineteen processing layers with: - two conv3-64 convolution layers; - followed by a maxpool pooling layer; - followed by two convolution layers conv3-128; - followed by a maxpool pooling layer; - followed by four convolution layers conv3-256; - followed by a maxpool pooling layer; - followed by four convolution layers conv3-512; - followed by a maxpool pooling layer; - followed by four convolution layers conv3-512; - followed by a maxpool pooling layer; - followed by three interconnection layers FC-4096, FC-4096 and FC-1000 and a Soft-max identification layer.
[0042] To conclude, there are a large number of neural network standards resulting from many years of research which have made it possible to arrive at particularly relevant configurations, in particular for carrying out identification functions.
[0043] However, the tests carried out on these different standardized configurations and in particular on the “VGG 16” and “VGG 19” configurations, did not allow satisfactory identification rates to be achieved for identifying the rooms, in particular for distinguishing a bathroom from a kitchen in a house.
[0044] The technical problem of the present invention is to obtain a convolutional neural network making it possible to more efficiently detect a type of room in a building. Statement of the invention
[0045] The invention proposes to address this technical problem by means of an architecture close to the “VGG 19” configuration but in which two convolution layers are removed. This specific architecture surprisingly makes it possible to obtain more relevant results than the “VGG 16” and “VGG 19” configurations.
[0046] Indeed, while it is not easy to understand why a particular architecture has better detection rates than another architecture, it is clear that the more the depth of a network increases, that is to say the greater the number of processing layers, the more the neural network is capable of identifying precise elements on the basis of training images. However, it is clear from this invention that the effective identification of a part requires the identification of a large number of salient objects but that the precise search for elements goes against the effective detection capacity of the neural network.
[0047] The present invention therefore resides in the detection of an effective compromise for the detection of parts between two known architectures, the “VGG 16” and “VGG 19” configurations.
[0048] Thus, the invention therefore relates to a method for identifying a room of a building by means of a convolution neural network making it possible to determine whether an input image corresponds to a predetermined type of room by means of a learning phase during which several images of each predetermined type of room are used, the convolution neural network comprising the following layers arranged one after the other to process the input image: - two convolution layers; - a grouping layer; - two convolution layers; - a grouping layer; - four convolution layers; - a grouping layer; - four convolution layers; - a grouping layer; - two convolution layers; - a grouping layer; - three layers of interconnection; and - an identification layer.
[0049] Preferably, as for the convolution neural networks “VGG 16” and “VGG 19”, the convolution layers are made with filters of 3 pixels by 3 pixels, a displacement between two filters of one pixel, a resizing of type “SAMEpadding” and using activation functions of type “RELU”.
[0050] Preferably, the grouping layers also use 2 pixel by 2 pixel filters, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of type "VAL1D padding".
[0051] Preferably, the three interconnection layers are successively: - a fully connected interconnection layer comprising 4096 neurons, - a fully connected interconnection layer comprising 4096 neurons, - a fully connected interconnection layer comprising 1000 neurons.
[0052] As for the “VGG 16” and “VGG 19” convolution neural networks, the format of the input image preferably corresponds to an RGB image of 224 by 224 pixels.
[0053] With this particular architecture of the neural network, the invention makes it possible to obtain an efficient identification of a room. For example, with training of the neural network at 100 epochs, that is to say by using 100 times the training data to modify the weight of the neurons during the training phase of the neural network, it is possible to detect with a very high recognition rate a bedroom, a kitchen, a bathroom or a living room of a home.
[0054] Of course, the recognition rate of the neural network also depends on the quality of the training data provided to it. However, the volume of training data currently available to characterize a part is still low compared to the volume of data currently accessible to distinguish, for example, a pedestrian from a bicycle in all the identification applications currently used for the development of autonomous vehicles.
[0055] Thus, the lack of training data is also a hindrance to the performance of neural networks for identifying a room and the invention also proposes to use a procedural generation of rooms in a home to address this lack of learning data. In this procedural generation, a predefined part is preferentially obtained with the following steps: - definition of a surface of the part; - definition of a part shape from the predefined surface; - generation of the walls and ceiling of the room; - placement of different objects characteristic of the room in a coherent manner; - allocation of textures to the floor, walls, and where applicable to the ceiling and / or objects in the room; - placing at least one camera and at least one light in the room; and - generating at least one realistic photographic rendering of the room.
[0056] As for the first step of defining a surface of the part, it is possible to randomly select a surface: - between 12 and 20 m2 for a bedroom; - between 8 and 15 m2 for a bathroom; - between 30 and 50 m2 for a living room; and - between 12 and 20 m2 for a kitchen.
[0057] As regards the second step of defining a shape of the part from the predefined surface, it is possible to implement the following steps: - random selection of a shape from a set of predefined shapes; and - applying a homothety to the selected shape so that the surface of the selected shape corresponds to the predefined surface.
[0058] As for the third step of generating the walls and ceiling of the room, it is possible to randomly select at least one height between 2.3 and 4 m, to raise walls to the selected height and to generate the shape of the ceiling according to the position of the upper end of the walls.
[0059] As regards the fourth step of placing different objects characteristic of the room, a catalog of objects each associated with one or more labels can be used.
[0060] Of course, other implementations of these steps are possible without changing the invention. Summary description of the figures
[0061] The manner of carrying out the invention as well as the advantages which result therefrom will emerge clearly from the following embodiments, given for informational but non-limiting purposes, with the support of the figures in which:
[0062] [Fig.l] is a table illustrating six standardized structures of state-of-the-art VGG-type neural networks;
[0063] [Fig.2] is a schematic representation of the layers of a neural network of convolution used in a preferred embodiment of the identification method of a part according to the invention
[0064] [Fig.3] is a schematic representation of the steps of a learning phase of the method for identifying a part according to a particular embodiment of the invention;
[0065] [Fig.4] is a photographic rendering of a living room produced by the procedural generation used in the method of [Fig.3];
[0066] [Fig.5] is a photographic rendering of a kitchen made by the procedural generation used in the method of [Fig.3];
[0067] [Fig.6] is a photographic rendering of a bathroom produced by the procedural generation used in the method of [Fig.3]; and
[0068] [Fig.7] is a photographic rendering of a room made by the procedural generation used in the method of [Fig.3]. Detailed description of the invention
[0069] The present invention relates to a method for identifying a room, i.e. a space inside a building delimited by at least one wall and / or at least one partition, for example a room in an apartment, a house or an industrial premises.
[0070] The method for identifying a part makes it possible to determine whether an input image corresponds to a predetermined type of part, the convolution neural network comprising the following layers arranged one after the other to process the input image: - two convolution layers, each made for example with 3 pixel by 3 pixel filters, a displacement between two filters of one pixel, a “SAME padding” type resizing and using “RELU” type activation functions, each layer comprising 64 neurons; - a grouping layer, created for example with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VALID padding” type; - two convolution layers, each made for example with 3 pixel by 3 pixel filters, a displacement between two filters of one pixel, a “SAME padding” type resizing and using “RELU” type activation functions, each layer comprising 128 neurons; - a grouping layer, created for example with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VALID padding” type; - four convolution layers, each made for example with filters of 3 pixels by 3 pixels, a displacement between two filters of one pixel, a resizing of type “SAME padding” and using activation functions of type “RELU”, each layer having 256 neurons; - a grouping layer, created for example with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VALID padding” type; - four convolution layers, each made for example with 3 pixel by 3 pixel filters, a displacement between two filters of one pixel, a “SAME padding” type resizing and using “RELU” type activation functions, each layer comprising 512 neurons; - a grouping layer, created for example with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VALID padding” type; - two convolution layers, each made for example with 3 pixel by 3 pixel filters, a displacement between two filters of one pixel, a “SAME padding” type resizing and using “RELU” type activation functions, each layer comprising 512 neurons; - a grouping layer, created for example with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VALID padding” type; - three interconnection layers, implemented for example with fully connected layers, the consecutive layers comprising respectively 4096, 4096 and 1000 neurons; and - an identification layer.
[0071] A preferred embodiment of the neural network used in the invention is illustrated in [Fig.2].
[0072] The input image is preferably an RGB image of 224 by 224 pixels. The method may comprise a conversion and / or resizing step to process an image which does not correspond to these parameters.
[0073] In order to train the neural network, the method according to the invention comprises a learning phase, during which parts are produced by procedural generation. These parts are generated to fit into a well-defined category, for example kitchen or bathroom, and are provided to the neural network to train it to recognize parts of the same category.
[0074] In procedural generation, we want to introduce as much variety as possible in the generation so that artificial intelligence does not make false associations between certain factors that are too recurrent when they are not necessarily so in reality.
[0075] Procedural generation allows for the generation of realistic photos of rooms from a number of categories, for example four categories: living room, kitchen, bedroom, bathroom.
[0076] The procedural generation of parts, an example of which is illustrated in [Fig.3], comprises the following steps: -definition 10 of a surface of the part; - definition 11 of a shape of the part from the predefined surface; - generation 12 of the walls and ceiling of the room; - placement of 13 different objects characteristic of the room in a coherent manner; - allocation of textures to the floor, walls, and where applicable to the ceiling and / or objects in the room; - placement of at least one camera and at least one lighting in the room; and - generation 16 of at least one realistic photographic rendering of the part.
[0077] The procedural generation is for example carried out by means of an algorithm written in Java language, using a catalog of objects listed in a file in XML format, and used as a plugin for the Sweet Home 3D® software.
[0078] During the surface definition step 10, the size of the part can be variable depending on the selected category. The size can be chosen randomly from common dimensions of real parts, for example: - bathroom: between 8 and 15 m2, - bedroom: between 12 and 20m2, - kitchen: between 12 and 20m2, - living room: between 30 and 50m2.
[0079] Once the surface is defined, the shape definition step 11 is implemented. In a first step, a number of walls can be chosen randomly, for example between four and six walls.
[0080] If the room has four walls, the following method can be implemented: - random generation of the width of the part, the width being included in an interval which may depend on the type of part and / or the surface; - calculating the length to obtain a rectangular shape.
[0081] If the room has five walls, to simplify the process, it is possible to limit the shape to a shape comprising four walls forming a rectangle, one corner of which is cut by a fifth wall. The following process can be implemented: - random generation of the width of the part, the width being included in an interval which may depend on the type of part and / or the surface; - random generation of the length of the corner wall, the length being within an interval, for example between 1 / 3 and 2 / 3 of the width of the room; - calculation of the length of the other walls, taking into account the width of the room and the length of the corner wall, the angle between the corner wall and the adjacent walls being for example fixed at 135°.
[0082] If the room has six walls, to simplify the process, it is possible to limit the shape to a shape comprising four walls forming a rectangle, two corners of which are cut by a fifth wall and a sixth wall. The following process can be implemented: - random generation of the width of the part, the width being included in an interval which may depend on the type of part and / or the surface; - random generation of the lengths of the first and second corner walls, the length being within an interval, for example between 1 / 3 and 2 / 3 of the width of the room; - calculation of the length of the other walls, taking into account the width of the room and the lengths of the corner wall, the angle between the corner wall and the adjacent walls being for example fixed at 135°; - random choice of the shape of the room: the two corner walls can be adjacent to the same wall or not.
[0083] For the step 13 of placing different objects characteristic of the room in a coherent manner, different types of objects grouped by types are used according to the category of the room, each type having a given function, for example: - bathroom: door, wash, toilet, sink, mirror, storage, - bedroom: door, bed, storage, floor, side, - kitchen: door, oven, hob, extractor hood, refrigerator, table, microwave oven, - living room: door, table, floor.
[0084] A catalog comprising a variety of objects of all types is used, each object having one or more labels corresponding to a type.
[0085] Some objects can cover two types at once, for example a cooker is in the type "oven" and "hob".
[0086] Certain types of objects may be mandatory and / or limited to a certain number per room. The minimum and maximum quantities of objects that the room can contain may be set arbitrarily, for example between one third and two thirds of its total surface area.
[0087] A tag can be associated with each type of object, so that the algorithm knows whether a given type of object is optional or mandatory. For example, since a door is mandatory in all rooms, we can associate a "mandatory" tag with the door type, or not associate an "optional" tag with it, depending on the tagging mode chosen.
[0088] Another "unique" tag can be used to mark types of objects that cannot be found more than once in the same room. An oven, for example, in a kitchen, or a bed in a bedroom.
[0089] Objects are placed randomly in the room, except for some that have specific placement rules, for example, a mandatory space may be provided around a table, or around a bed except for the headboard, or in front of a door or a window. Besides special cases, most objects can be placed along a randomly chosen wall, taking into account the placement of other objects so as not to embed them into each other. Thus, when an object and a wall have been randomly chosen, the algorithm can look at what furniture is already arranged along this wall, and check whether the free segments of the wall can accommodate the new piece of furniture to be placed.
[0090] It may also be taken into account in the placement to arrange close to each other the objects usually close, with their usual orientation, such as kitchen furniture, a bed and bedside tables, a table and chairs, a coffee table in front of a sofa or even armchairs in front of a television or a fireplace.
[0091] Regarding openings, i.e. doors and windows, the algorithm can prevent a window and an interior door from being placed on the same wall, for example. As with other objects, the number of openings can be generated randomly depending on the type of room and the surface area.
[0092] The object catalog can be a file in XML format, including for each object its type(s), and for each type optionally one or more tags.
[0093] During the step 14 of assigning textures to the floor, walls, and where appropriate the ceiling and / or objects in the room, a texture catalog may be used. Random textures may be used without taking into account the room category, so that the identification algorithm does not give too much importance to the material of the floor and walls when attempting to identify the room type.
[0094] The step 14 of assigning textures can be carried out before the step 13 of placing the objects.
[0095] Steps 10 to 14 provide a 3D model of a part.
[0096] During the step 15 of placing at least one camera and at least one light in the room, the Blender® software can be used. In order to generate a random placement of the lights and the camera in the room, a Python script controlling this software can be used.
[0097] The “Cycles” rendering engine of the Blender® software can then be used to manipulate the camera(s) and the lighting(s) to have several angles and in different lighting configurations (start of day or end of day for example).
[0098] Step 16 of generating a realistic photographic rendering can then be implemented.
[0099] Figures 4 to 7 illustrate examples of images obtained by procedural generation, at the end of generation step 16. [Fig.4] represents a living room, [Fig.5] a kitchen, [Fig.6] a bathroom and [Fig.7] a bedroom.
[0100] Realistic photographic renderings can then be generated in number in order to constitute a database of photos of various parts belonging to the desired categories. This database can then be used for training the neural network. The detection of the type of part will then be all the more efficient if the database used for training contains a large number of photos, and if these photos illustrate the diversity of parts within each category.
Claims
Claims
1. A method for identifying a room of a building by means of a convolutional neural network for determining whether an input image corresponds to a predetermined room type by means of a training phase during which several images of each predetermined room type are used, the convolutional neural network comprising the following layers arranged one after the other to process the input image: - two convolution layers (conv3-64); - a pooling layer (maxpool); - two convolution layers (conv3-128); - a pooling layer (maxpool); - four convolution layers (conv3-256); - a pooling layer (maxpool); - four convolution layers (conv3-512); - a pooling layer (maxpool); - two convolution layers (conv3-512); - a pooling layer (maxpool); - three interconnection layers (FC-4096, FC-1000);and - an identification layer (Soft-max).;
2. Method for identifying a part according to claim 1, in which the convolution layers (conv3-64, conv3-128, conv3-256, conv3-512) are made with filters of 3 pixels by 3 pixels, a displacement between two filters of one pixel, a resizing of type "SAME padding" and using activation functions of type "RELU".
3. Method for identifying a part according to claim 1 or 2, in which the grouping layers (maxpool) are produced with filters of 2 pixels by 2 pixels, a function for calculating the grouping pixel using the maximum, a displacement between two filters of two pixels and a resizing of the “VAL1D padding” type.
4. Method for identifying a part according to one of claims 1 to 3, in which the three interconnection layers are successively: - a fully connected interconnection layer (FC-4096) comprising 4096 neurons, - a fully connected interconnection layer (FC-4096) comprising 4096 neurons, - a fully connected interconnection layer (FC-1000) comprising 1000 neurons.
4. Method for identifying a part according to one of claims 1 to 3, in which the input image corresponds to an RGB image of 224 by 224 pixels.
5. Method for identifying a room according to one of claims 1 to 4, in which the learning phase is carried out by procedural generation of rooms with, for each type of room, the following steps: - definition (10) of a surface of the room; - definition (11) of a shape of the room from the predefined surface; - generation (12) of the walls and the ceiling of the room; - placement (13) of different objects characteristic of the room in a coherent manner; - allocation (14) of textures to the floor, the walls, and where appropriate to the ceiling and / or to the objects of the room; - placement (15) of at least one camera and at least one lighting in the room; and - generation (16) of at least one realistic photographic rendering of the room.
6. Method for identifying a room according to claim 5, in which the step of defining (10) a surface area of the room consists of randomly selecting a surface area: - between 12 and 20 m2 for a bedroom; - between 8 and 15 m2 for a bathroom; - between 30 and 50 m2 for a living room; and - between 12 and 20 m2 for a kitchen.
7. Method for identifying a part according to claim 5 or 6, wherein the step of defining (11) a shape of the part from the predefined surface implements the following steps: - random selection of a shape from a set of predefined shapes; and - application of a homothety to the selected shape so that the surface of the selected shape corresponds to the predefined surface.
8. Method for identifying a room according to one of claims 5 to 7, in which the step of generating (12) the walls and the ceiling of the room consists of randomly selecting at least one height
9. between 2.3 and 4 m, to raise walls to the selected height and generate the ceiling shape based on the position of the upper end of the walls. Method for identifying a part according to one of claims 5 to 8, in which the step of placing (13) different objects characteristic of the part uses a catalog of objects each associated with one or more labels.