Computer-implemented object recognition method
The computer-implemented object recognition method addresses the challenge of identifying objects of varying sizes by using computer-generated data sets and recognizing size ratios, resulting in improved accuracy and reliability of object recognition.
Patent Information
- Application Number
- PCT/EP2024/086260
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-15
- Filing Date
- 2024-12-13
- Publication Date
- 2025-06-19
AI Technical Summary
Existing object recognition methods struggle with accurately identifying object bodies in real scenes, particularly when objects of the same type differ only in geometric size, leading to high error rates and reduced reliability.
A computer-implemented object recognition method that utilizes high-quality computer-generated image data sets and recognizes size ratios to unambiguously identify object models, even when they differ in geometric size, by decoupling scale from object recognition and using a neural network to predict necessary transformations.
This approach significantly reduces confusion errors and increases the validity of object recognition, enabling precise identification of object bodies even with wide size variations, thus enhancing the reliability of object recognition in industrial and other sensitive applications.
Smart Images

Figure EP2024086260_19062025_PF_FP_ABST
Abstract
Description
[0001] Computer-implemented object recognition method
[0002] The invention relates to a computer-implemented object recognition method for identifying object bodies in a real scene.
[0003] Object recognition methods of this type are known in different forms.
[0004] For object recognition, it is known to use pattern / object pattern recognition methods to infer components in a real scene. The real scene is captured for computer-aided or computer-implemented processing using an image capture device in order to analyze the relevant image data set for object recognition. Such methods are known, for example, under the term "convolutional neural networks."
[0005] Digital cameras, for example and in particular, are commonly used as image capture devices to capture a real scene as a two-dimensional image and to provide it as a digital image data set for further computer-aided or computer-implemented processing.
[0006] For object recognition in a real scene based on images, the use of filter methods is known for the relevant methods, for example to filter out an object pattern, such as an object edge, line, texture and the like, based on contrast changes or jumps in the image contained in the digital image data set and to infer the object body on the basis of this.
[0007] Based on the recognized object patterns, a Kl module is able to recognize an object body by comparing object patterns with object bodies already recorded as object models and finally to identify it in the image.
[0008] For this purpose, in a learning or training process of a respective method, real scenes are first recorded and labeled based on the known object bodies contained therein. The metadata then stores the data with metadata. Based on this metadata, such a method is capable of recognizing and identifying these object bodies in any constellation of real scenes. The metadata includes, for example, and in particular, information about the object models of a scene arrangement, their relative arrangement, and other data relating to the acquisition properties, such as the acquisition perspective.
[0009] The training process is carried out, among other things, with the help of interactions with a trainer or operator who manually marks the objects and uses their input to train the procedure.
[0010] The aim of computer-implemented object recognition processes is to reduce manual interventions, so that these processes are increasingly being improved with the involvement of so-called artificial intelligence (acronym: AI).
[0011] Therefore, computer-implemented object recognition methods for detecting and identifying object bodies in a real scene are known, which have the following detection steps:
[0012] • Providing an image data set of a real scene captured by an image capture device for evaluation by a Kl module,
[0013] • Analysis of the image data set by the Kl module and recognition of object patterns in the real scene contained therein,
[0014] • Identification of object bodies contained in the real scene based on the recognized object patterns by the Kl module.
[0015] A KL module is a computing system with a neural network used to recognize object patterns in the real scene. A neural network is a technology for distributed content storage in a network whose structure is modeled on that of the human brain. Content is collected and weighted against each other according to its relevance to a query, generating a result based on the weighting.
[0016] The term neural network is also known as "artificial neural network" or "simulated neural networks." The same applies to the term artificial intelligence or its acronym Kl.
[0017] To successfully use a Kl module with a neural network, training with training data is first necessary to train the neural network. The training of a Kl module is complete when the reliability of the result reaches a sufficiently high level and consistently does not fall below this level. This level depends on the specific application. Within the scope of the invention, an image data set can be processed or manipulated using computer-aided / computer-implemented processing, particularly a digital image data set. Accordingly, the term "image data set" is used for short within the scope of the invention.
[0018] The quality of an image dataset and its metadata used to train a CL module is crucial for the subsequent performance and generalization capability of object recognition. For example, variations in environmental conditions such as lighting and background properties must be taken into account. Furthermore, the image dataset is usually manually enriched with metadata, which is both time-consuming and error-prone.
[0019] In many cases, the effort required to create suitable training data is therefore too high, which is why the known object recognition methods of the type in question often still have a high error rate with regard to object identification, so that their reliability is still insufficient for many application areas.
[0020] Maintaining a low error rate in object recognition is particularly important in sensitive application areas, including security-related, medical and industrial applications.
[0021] Against this background, methods for creating synthetic training data to reduce manual effort are known, for example, from WO 2023 / 014369 A1 and US 2020 / 302241 A1. Furthermore, US 2020 / 342652 A1 describes a system for generating synthetic images to create them with different perspectives and use them for a learning or training process of an AI module. To do this, a synthetic scene is first arranged by combining different object models, background images, exposures, and camera settings.
[0022] The known object recognition methods aim to recognize an object type without making unambiguous inferences about the specific object body, and to differentiate them from one another in such a way that an object body of one object type recognized in a scene arrangement can be distinguished from another object body of the same object type. The known methods reach their recognition limits, particularly in cases where a scene arrangement contains object bodies that differ from similar object bodies of the same object type only in their geometric size.
[0023] Against this background, the object of the invention is to provide a computer-implemented object recognition method that reduces the error rate for the recognition of object bodies in a real scene.
[0024] A computer-implemented object recognition method is also referred to as an object recognition method.
[0025] The invention departs from the approach of increasing the recognition rate by training computer-implemented object recognition methods with real scenes.
[0026] The invention also moves away from the approach of changing the algorithms with which a Kl module internally processes processes.
[0027] Rather, the invention solves its problem by recognizing a scale in a scene arrangement and, based on this, enabling the unambiguous recognition of object models contained therein that are similar or identical to object models in their design, but differ from one another in their actual geometric size. The invention thus avoids the error of incorrectly recognizing an object body based on its design. To this end, the invention utilizes the recognition of size ratios and a geometric scaling with which the object models are reproduced or contained in the digital representation.
[0028] The invention thus closes the gap for application in industrial fields where object bodies must be clearly identified in order to precisely initiate and control further process steps in a product value chain on this basis. For example, the invention now makes it possible to precisely initiate ordering and production processes.
[0029] The invention succeeds in clearly recognizing object bodies in a scene arrangement, even in the case of object bodies with a wide range of sizes within graphical representations of a scene arrangement. The invention thus reduces confusion errors and thus increases the validity of object recognition. The invention achieves this objective by training the AI module with its neural network using high-quality computer-generated image data sets and taking the respective size ratios in the digital representations into account for object recognition.
[0030] The invention is based on the approach of using at least high-quality object model data sets and combining them with background data sets as synthetically prefabricated components to form a scene arrangement. This scene arrangement serves to enable precise training of the AI module. Therefore, according to the invention, a respective scene arrangement is, for example and in particular, such that at least the object models contained therein digitally recreate or replicate components of a possible or actual real scene. To this end, the invention uses the creation of an image pool with, in particular, complex or high-resolution images. These images are therefore synthetic images.
[0031] High-quality object data sets are in particular those that digitally recreate or reproduce object bodies relevant for object recognition using object models, particularly true to scale.
[0032] This has the advantage of positively influencing the learning curve for object recognition within any real scene or synthetic scene arrangement. This increases the reliability of object recognition with minimal time expenditure. Furthermore, user interventions can be minimized, thus also reducing the effort required to train the AI module.
[0033] To this end, the invention provides that the Kl module for recognizing and identifying objects is trained by the following iteration steps: i. Providing a supply of object data sets for object models, ii. Providing a supply of background data sets for background models, iii. Providing a plurality of image data sets for a digital scene arrangement, wherein the image data sets are created by means of a scene creation module, iv. Analyzing the respective image data set and recognizing object patterns in the scene arrangement contained therein, v. Identifying the object model(s) contained in the scene arrangement, wherein a. at least one fragment of the object patterns and / or a background pattern is detected in the scene arrangement as a reference for determining a scale, b. the scale of the graphical representation is specified c.the respective object pattern contained in the scene arrangement is recognized based on the scale, i.e. the respective object model is determined based on the respective object pattern, vi. Verifying the accuracy of identifying object bodies based on the metadata of the digital scene arrangement, vii. Repeating the iteration steps if the verification results in insufficient accuracy.
[0034] The occurrence of confusion errors is primarily due to the fact that, in order to account for as many variations as possible that occur in real scenes, the distance between the virtual camera and the object models is varied during the creation of training data. However, this means that standard object recognition methods are no longer able to clearly identify size variations of an object type.
[0035] One possible solution would be to keep the camera distance constant in the virtual scene and thus provide a benchmark for taking images of a real scene, which would, however, severely limit the areas of application of object recognition.
[0036] Furthermore, the error rate in object recognition would only fall to an acceptable level if the image data sets to be analyzed for a captured real scene were also recorded with comparable settings, or if these settings were known for the evaluation of the image data set. However, these settings are neither necessarily valid nor complete. Therefore, the invention uses the captured real scenes without metadata for the camera settings. This advantageously enables the invention to dispense with specifications or metadata for image acquisition and to analyze a broad, in particular unlimited, spectrum of captured real scenes.
[0037] The invention solves this problem by decoupling the scale from the actual object recognition. Accordingly, the variation in camera distance is limited when creating the synthetic training dataset for object recognition, thus avoiding confusion errors if the input images adhere to the thereby defined reference scale.
[0038] To achieve this, the AI module is supplemented with another neural network, which is trained with strong scale variations to predict the necessary transformation to reconstruct the reference scale. This allows the recording of a real scene to be cleverly preprocessed, significantly improving subsequent object recognition.
[0039] In one method option, the invention takes into account the determination of a scale, for example, and in particular, by identifying potential reference objects in the pool of data sets for object models based on the data sets. Such a reference object is present, in particular, if no other data set with a design duplicate exists for the data set in question, which differs from the first-mentioned data set only by a geometric scaling.
[0040] Thus, an object recognition method according to the invention can use such a constellation to detect a reference in a digital representation and thus to establish a scale.
[0041] For this purpose, the object pattern of the reference can be compared with the geometric data stored in the relevant data set. Alternatively or additionally, it is possible within the scope of the invention to perform a first object recognition in order to compare the recognized objects with each other with regard to their sizes present in the digital representation, in order to infer an object reference for determining a scale based on the size comparison.
[0042] Furthermore, fragments, especially geometric fragments, can also serve as references in a digital representation and be used accordingly. For example, it is possible to detect at least one pattern in a background in order to use this as a reference for determining the scale. In this context, it is also possible to detect a reference based on geometric lines or arrangements in order to determine a scale in a digital representation based on geometric relationships of lines, angles, and the like.
[0043] Finally, the object model belonging to an object pattern is determined based on the scale, on the basis of which the respective object body to which the respective object model is assigned is identified.
[0044] Within the scope of the invention, the term "arranging" encompasses the possibility of arranging object and background models in a scene arrangement. Therefore, within the scope of the invention, "arranging" encompasses the arrangement of one or more object models to form a background. The object(s) are each arranged at a distance and an orientation relative to the background. When using multiple objects, the respective distances and orientations can be different from each other and from the background. Furthermore, the adjustment of the distances and orientations can follow a fixed pattern or be done arbitrarily.
[0045] The invention also considers the options for changing or varying components of a background model and an object model, for example, and in particular, scaling them and aligning them in different positions relative to one another or to the background. Furthermore, arranging within the scope of the invention also considers the option of creating variants, for example, in the representation of the object and the background, or in the generation of objects and backgrounds. Thus, the invention encompasses the arrangement of the object and background models by arranging various components to form a scene arrangement.
[0046] Within the scope of the invention, metadata also includes tags for objects or a recording or a scene arrangement. The metadata is used to create so-called labels for objects or arrangements within the scope of machine learning or artificial intelligence methods. The labels for the Kl module are derived from the image data or metadata and can in particular also include metadata. The method of using labels is also referred to as so-called supervised machine learning within the scope of Kl. Within the scope of the invention, the term real scene takes into account an image section from a scene that can originate from a real scene. According to the invention, the term real scenes also includes images that are displayed on an output device, for example a screen, and can be captured by an image capture device.
[0047] Within the scope of the invention, image capture devices include devices and devices designed and configured to capture real scenes for computer-aided or computer-implemented processing. These include, for example, and in particular, digital cameras that create corresponding recordings and provide them as computer-processable images. The images can also be moving images, such as videos or animations.
[0048] Within the scope of the invention, a digital representation is a captured scene arrangement that can be digitally processed and stored in an image data set. According to the invention, it comprises an image of the scene arrangement captured from one or more capture perspectives. Capture occurs, for example, and in particular, using a virtual camera that captures the scene arrangement while positioned and aligned. The section of the scene arrangement captured by the virtual camera is adjustable, and in particular, the object or objects can also be adjusted in relation to the background.
[0049] Metadata for a digital scene arrangement can, for example, and in particular, relate to information about the position, orientation, and size of an object contained in the scene arrangement. This enables the invention, for example, and in particular, to verify the detected scale in a digital representation and, in the event of deviations, to further train the AI module.
[0050] Within the scope of the invention, the object data sets and the object models contained therein are those that can be processed or edited using computer-aided or computer-implemented methods. Therefore, these models and data sets are available in a digital format.
[0051] The same applies to the background data sets and the background models they contain. The background models serve as the background or environmental representation of a scene, in which one or more object models are placed. Therefore, background data sets are also available in a digital format.
[0052] In order to take into account different light and contrast effects, the invention includes the structure of the background being at least a part of the light source and therefore the scene arrangement being able to be influenced by different backgrounds.
[0053] The object data sets and the background data set are combined into a scene arrangement using a scene creation module. Within the scope of the invention, a scene creation module is a computer-implemented image or model processing application that allows the creation of a digital image data set by arranging object and background models.
[0054] Such scene creation modules are commonly known, for example, as image processing and modeling programs, applications, or tools. They offer the ability to arrange and modify object and background models and ultimately save them as an image dataset.
[0055] Within the scope of the invention, an image data set can be designed in a variety of ways, for example and in particular as a 2D representation or image in an image format as well as as a 3D representation in a 3D image format.
[0056] Furthermore, the invention includes the possibility of designing the image data sets for recording moving images, such as a sequence of images, animations or videos, and processing them accordingly.
[0057] Accordingly, within the scope of the invention, not only static representations but also dynamic digital representations of real scenes and scene arrangements can be used for object recognition. Accordingly, corresponding image capture devices that provide static and / or dynamic representations in a digital image data set can also be used.
[0058] Within the scope of the invention, the object models serve as digital representations of possible or existing object bodies in a real scene. According to the invention, they are digitally modeled or reproduced accordingly. The same applies to the background models, which are also digitally modeled or reproduced according to possible or existing environments. According to the invention, the image data sets are created by means of a scene creation module through the following steps: a. Reading in a, in particular randomly selected, background data set, b. Reading in a, in particular randomly selected, object data set or a plurality of, in particular randomly selected, object data sets, c. Arranging the object model of the object data set or the object models of the object data sets and the background model of the background data set to form a scene arrangement, d. Capturing the scene arrangement in a digital representation, e.Storing the digital representation in an image dataset and storing metadata for the digital representation of the scene arrangement.
[0059] This has the advantage that scene arrangements can be created with a spectrum of variations in the components of a particular background, the arrangement, the orientation as well as the representation of object models.
[0060] Due to the diverse variations, the invention succeeds in training the Kl module with adverse constellations, so that object recognition is accordingly increased.
[0061] Within the scope of the invention, an object model is a digital representation of a real or potentially real object body. Within the scope of the invention, the object model is preferably defined three-dimensionally.
[0062] The creation of a scene arrangement can be based on a predetermined selection, for example by following a pattern predetermined in a list or script.
[0063] According to the invention, the data sets can be selected from the pool in the order in which they are contained therein. However, in order to increase the variation possibilities for obtaining a multitude of scene arrangements, the invention includes the possibility of selecting the object model(s) randomly, in particular stochastically.
[0064] Within the scope of the invention, it is taken into account that an object model comprises a plurality of components, in particular, it can consist of a plurality of elements as well as shapes. The same applies to a background model contained in a background data set, which, according to the invention, represents the environment of an object.
[0065] To increase reliability, the invention provides that parameters with which the scene arrangement is captured in a representation are varied and stored with the effects of the variations in an image data set.
[0066] To this end, the invention comprises the ability to change the parameters with which a representation is stored in an image dataset. These parameters include, for example, and in particular, perspective settings, an orientation and a detection range of a virtual camera, exposure or illumination values, and color and contrast spectra, as are known from image recordings. These parameters can preferably be stored as metadata for a scene arrangement or an image dataset.
[0067] The invention provides that the iteration steps of a computer-implemented object recognition method are designed in such a way that the Kl module is trained automatically.
[0068] This offers the advantage that the individual steps are carried out automatically, so that human intervention is at least reduced to a minimum or even eliminated. This means that the invention enables training of the AI module in a wide variety of ways and with minimal expenditure of time and resources.
[0069] High data quality and integrity are required in information-sensitive application areas. These include industrial tasks and applications, such as those found in electrical engineering. In the industrial sector, high availability and reliability of industrial systems are required, which places correspondingly high demands on the operational reliability of electrical installations and their maintenance and repair.
[0070] The secure object detection of electrical components in a control cabinet presents employees as well as technical detection systems with significant challenges regarding the reliable and unambiguous detection and identification of object bodies.
[0071] This is due, among other things, to the multitude of components arranged in a control cabinet. Therefore, a fully equipped and operational control cabinet often contains numerous optical interference elements, which, while useful for operation, make object detection difficult.
[0072] These include, for example and in particular, marking labels for identifying components as well as connecting cables or electrical jumpers, which are sometimes found in a disorganized or non-schematic manner in a control cabinet.
[0073] A computer-implemented object recognition method according to the invention is particularly configured and designed to ensure rapid and reliable identification of objects. Against this background, the invention encompasses the use of a computer-implemented object recognition method according to the invention for electrical engineering applications or installations for detecting object bodies that are electrical components or assemblies, such as those found, for example, and in particular, within a real scene as components of a control cabinet. These include the aforementioned, such as terminal blocks, switching and measuring devices, electrical supply and conversion devices, as well as electrical distributors, connecting cables, and jumpers.The above list is not exhaustive, but represents examples of typical components of a real scene for an electrical engineering application.
[0074] Therefore, the invention considers that the object models comprise, for example, and in particular, electrical components or assemblies, such as those preferably mentioned above. Thus, the invention can preferably be used to depict electrical installations using scene arrangements and thus to recognize electrical object bodies in corresponding real scenes, as mentioned above in part. Thus, the real scenes within the scope of the invention can in particular be those of electrical installations.
[0075] Within the scope of the invention, it has been shown that generating image data sets using true-to-scale and detailed object models results in an increase in the reliability of object recognition compared to those based on photos of real scenes. This is primarily due to the inventive advantage that any constellations can be arranged that cannot be created in reality or can only be created with great time and expense. Furthermore, according to the invention, the effort required to create scene arrangements is significantly reduced compared to an arrangement for a real scene. Furthermore, the degree of freedom is increased and the adaptation effort required to create a variety of variants for scene arrangements is reduced. This enables fast and precise training of a Kl module, which in particular has the advantage of sharpening the Kl module's recognition accuracy in the event of errors.
[0076] Therefore, an advantageous development of the invention provides that the object model is contained in the object data set in a vector-based data format, in particular a CAD data format.
[0077] Accordingly, the object model contained therein is a vector-based object model or a CAD object model. This results in the advantage that the object models have a high degree of representational accuracy and are easily editable and adaptable for the respective application. Therefore, an object model is, in particular, a three-dimensional, particularly true-to-shape representation of an object body.
[0078] Furthermore, this results in the advantage that these object models or object data sets can already be used from other applications or company areas, for example and in particular product development, and thus the up-to-dateness of the object models or data sets is ensured.
[0079] This has the advantage that a wide range of object models is available, covering the entire product portfolio. Furthermore, the effort required to create the object models is eliminated, as these have already been created for other purposes, thus saving the time and resources dedicated to an object recognition method according to the invention.
[0080] Furthermore, this ensures that the object models include all adaptations and, for example, correspond, in particular, to the sales products. Furthermore, it ensures that the data and models are sufficiently up-to-date and eliminates the additional effort required for data maintenance. Another advantage is that capturing real objects using a digital camera is no longer necessary. This also eliminates the effort required to capture real objects from different perspectives. At the same time, it eliminates the need for a suitably set-up background against which the real object is photographed or digitally captured.
[0081] Within the scope of the invention, it has therefore been shown that databases already used operationally for other purposes in a company can be used to provide a stock in an object recognition method according to the invention.
[0082] Against this background, a further advantageous development of the invention provides that the stock of object data records comprises a spectrum of differently designed object models, in particular the spectrum of products of a product portfolio, preferably the object data records originate from a product database.
[0083] One advantage that is evident in industrial companies is that an object recognition method according to the invention can be coupled with internal databases as well as data warehouse systems or information systems, which enables mutual usability of the data and object recognition for different company divisions.
[0084] Within the scope of the invention, it has been shown that in many applications, an external representation is sufficient for object recognition and identification. In this regard, a further advantageous development of the invention provides that the object data set contains the object model in a representation with reduced external surfaces and / or edges.
[0085] This offers the advantage of simplified and resource-reduced inventory management. Furthermore, the reduced display reduces the computer support requirements and the time required for object recognition and identification.
[0086] This promotes corresponding object recognition in real time, as is taken into account in a further advantageous development of the invention by the detection steps being configured and designed such that identification occurs with a low latency, particularly in real time. This advantageously leads to high acceptance of an object recognition method according to the invention. According to the invention, precise object recognition in real time is achieved in particular by targeted training of the AI module, in that the system can be trained using numerous arrangements as well as a multitude of constellations and variants in a very short period of time compared to conventional methods.
[0087] For this purpose, the invention provides in an advantageous further development that the parameters for capturing a scene arrangement are set, in particular randomly or stochastically, before the representation is stored in an image data set.
[0088] The parameters can be set based on routines, such as a script, by specifying different values for the parameters. Furthermore, the invention also allows for the values to be set randomly, particularly stochastically. This allows for a broad spectrum of different scene arrangements, which results in different variants being taken into account during training. The invention thus makes it possible to provide different levels of difficulty for training the AI module, thereby promoting the reliability of object recognition.
[0089] Accordingly, an iteration step results that takes priority over storing the representation in an image dataset.
[0090] In a further advantageous development of the invention, it is provided that the object models or the background models are provided as parameterized models that can be adjusted for the creation of a scene arrangement.
[0091] This facilitates training of the Kl module in a surprisingly simple manner, so that within the scope of the invention it is advantageously achieved to provide a wide range of possible variations in order to increase the reliability of object recognition.
[0092] The change in parameter settings for variant creation can, in turn, be set based on routines, such as instruction lists or scripts, or also randomly or stochastically, so that the process of creating variants when creating object models and background models can be automated and thus implemented accordingly efficiently. For example, and in particular, parameters relate to geometric dimensions, colors, light values, as well as the arrangement of object models relative to one another or relative to a background model. For this purpose, the invention provides for providing object models and background models with parameters that can be variably adjusted accordingly.
[0093] Within the framework of the proposed computer implementation, such changes can be implemented in various ways, allowing the settings to be generated automatically. This provides both time and cost advantages for the invention.
[0094] The arrangement and alignment of the object models can be achieved by selecting a reference, so that the object models are arranged relative to this reference. According to the invention, the arrangement can be random, stochastical, or predefined. The reference can be a point, a line, or another structure that is arranged transparently or opaquely in the scene arrangement.
[0095] For the reality-oriented digital reproduction of ordered constellations, a further advantageous development of the invention provides that the arrangement of a plurality of object models to one another takes place at least according to one of the following options and therefore comprises at least one of the following characteristics:
[0096] • Row arrangement,
[0097] • Column arrangement,
[0098] • Matrix arrangement,
[0099] • equally spaced arrangement.
[0100] The order of the object models can be determined randomly, especially stochastically. Furthermore, it is possible to select an arrangement and then select the object models and their order to create a scene arrangement randomly or stochastically.
[0101] The invention makes it possible to control and generate an arrangement using a routine, such as a combination list or a script. Accordingly, combination rules can be provided that control the relative arrangement of the object models. This can advantageously be used to specifically train the AI module and to mitigate or eliminate potential or identified weaknesses in object recognition through adapted background models. Particularly in application areas where a predefined order of object bodies prevails, this order should also be taken into account in the image data sets used to train the AI module.
[0102] Such fields of application are regularly found - as already mentioned - in the field of electrical and electronic systems, such as devices, equipment and circuit arrangements in control cabinets that have ordered constellations.
[0103] Therefore, the invention also encompasses an ordering according to which an object model serves as a reference object for arranging further object models relative to it. The arrangement of the objects relative to one another, as well as their selection from the pool, can be random, in particular stochastically, as already described above.
[0104] A background has a significant influence on the recognizability of an object. In addition to different patterns and textures, the color, light, and contrast effects on objects arranged in the background are also significant influencing factors that must be considered when training the AI module.
[0105] In a further advantageous development of the invention, the effect of the background on the recognizability of object bodies in a real scene is effectively taken into account when training the Kl module in that the arranging comprises projecting the background model contained in the background data set onto a curved, in particular concavely curved, preferably a hemispherical, surface.
[0106] This creates the advantage of creating a three-dimensional background that corresponds to a real scene. Accordingly, the environmental influences on object recognition can be simulated at least nearly comprehensively and completely and captured during training of the AI module. The invention thus differentiates itself further from previously known methods in favor of a synthetic, nearly life-identical simulation of detection situations in a real scene. In particular, the simulation of the exposure situation by taking into account realistic detection constellations and background interactions is surprisingly improved. The invention considers training of an AI with a scene arrangement from views that frequently occur in a real situation or are known to occur in practice, so that superfluous views can be excluded. This contributes to an improvement in training performance.In this respect, it has been shown, particularly in electrical engineering applications, that views in three orientations are sufficient, and additional views are unnecessary for training a control module. This allows the invention to increase the efficiency of the training process.
[0107] Within the scope of the invention, it has been shown that the exposure characteristics have a significant influence on object recognition. Therefore, it was determined that the influence of the background on the exposure of the object models in a scene arrangement must be taken into account in order to precisely train the AI module and improve object recognition.
[0108] Therefore, an advantageous development of the invention is directed toward generating the influence of the background model on the exposure of the respective object model and taking it into account accordingly when training the AI module. This also takes into account the reflection and shading effects of the background on the object model(s) or the interaction between the object model or the plurality of object models in a scene arrangement.
[0109] Against this background, a relevant development of the invention provides that the background data set is configured and designed in such a way that it effects exposure of the object model or the plurality of object models. An object model with such exposure adjustment is thus available for training the AI module, thereby resulting in further precision in object recognition.
[0110] According to the invention, this is achieved, for example and in particular, by simulating the light values of the background model based on the individual pixels and their influence on the exposure of an object or the majority of object models. This also includes capturing interactions with regard to the exposure values between the background model and the object model or object models and, for example and in particular, also taking into account reflection and shadow effects between them and taking into account the resulting light and contrast values for the respective object model when teaching the Kl module. According to the invention, the arrangement of the object models relative to the background model can be arbitrarily carried out.Within the scope of the invention, for example and in particular, the object model contained in an object data set can be arranged at a distance in the direction of a surface normal of the curved surface, preferably on a plane opposite the curved surface.
[0111] By reducing the degree of freedom of the arrangement of the object models in this way, it is possible to make the arrangement more realistic, so that in turn a class can be trained more specifically.
[0112] According to the invention, this can be enhanced by providing the background model contained in the background data set in the manner of a panoramic image. Panoramic images capture an environment or scene on a larger scale from multiple perspectives. A panoramic image can therefore contain a panoramic view in its entirety or in sections.
[0113] Various methods are known for creating a panoramic image. Such methods are already implemented in digital cameras, for example, so that, according to the invention, a real background can be captured or recorded in a panoramic image. Furthermore, applications, so-called stitching software, that serve this purpose are known.
[0114] This makes it possible to reduce the time required to train the AI module. Furthermore, the system requirements can be kept low. Furthermore, it is possible to quickly and comprehensively simulate backgrounds and easily vary them, allowing the AI module to be trained in a wide variety of ways and with great precision, something that previous approaches have not been able to achieve.
[0115] The invention includes the option that a background model is constructed using computer support and is available, for example, in the form of a three-dimensional digital image or as a CAD model.
[0116] To create an image dataset using a panoramic image, it has proven advantageous for the image dataset to have an image data format and to contain the background model as an image, in particular as a panoramic image, preferably a 360-degree panoramic image. Storing the image dataset in an image format is advantageous in that, according to the invention, the capture of a real scene can also be carried out in an image format that is particularly compatible with the image format of the image dataset used to train the AI module.
[0117] Based on this, the Kl module can be trained in a targeted manner so that compatibility problems or conversion-related differences between the image data sets are effectively avoided.
[0118] For this purpose, in a further advantageous development of the invention, at least the storage of the representation takes place in an image data set with an image data format, wherein this image data set contains in particular the background model as an image, in particular as a panoramic image, preferably as a 360-degree panoramic image.
[0119] Using a panoramic image for the background model offers the advantage that different spatial capture perspectives can be generated for training the AI module using a single representation. Only one scene arrangement can be used for different capture perspectives, each of which is stored in an image dataset and made available for the subsequent analysis and identification steps for training the AI module. This also applies to background models that do not use a panoramic image and instead use a different type of background model, such as a two-dimensional representation of a background.
[0120] Within the scope of the invention, it has been shown that at least the time required to create scene arrangements can be reduced by using a basic structure for a plurality of digital image data sets.
[0121] For this purpose, a further advantageous development of the invention provides that the capturing of the scene arrangement comprises a variation of a capturing perspective, wherein the variants of the capturing perspective are captured in a respective digital representation.
[0122] This creates a scene arrangement that is captured from different perspectives and stored in an image dataset using a respective representation. The metadata for the respective digital representation of the scene arrangement is also saved. This results in the advantage that training the AI module takes less time, as the variants do not have to be completely recreated. Furthermore, training can be targeted, for example, using the most frequently occurring perspectives in reality or based on insights into which perspective leads to a higher error rate in object detection and identification.
[0123] In the context of the invention, it has been shown that time savings for training the Kl module can be increased by using a scene arrangement as a basis for further variant formation.
[0124] For this purpose, a further advantageous development of the invention provides that the scene arrangement is recorded in such a way that a variation of at least one of the following aspects
[0125] • the background model,
[0126] • the selection of object models,
[0127] • the arrangement of the object models to each other,
[0128] • the light and contrast values of the scene arrangement are created and the variants of the scene arrangement are recorded in a respective digital representation.
[0129] As already explained analogously above, a scene arrangement is created and used to generate different variants of this scene arrangement based on it and to record the variants using a corresponding digital representation and then to save the respective digital representation in an image data set and to save metadata for the respective digital representation of the scene arrangement.
[0130] This makes it possible to simplify and accelerate the creation of various scene arrangements in a resource-neutral way. Furthermore, it makes it possible to train the AI module according to a scheme, ensuring successful training through routine processes.
[0131] The background model can be a photorealistic reproduction or a real background, which is captured by an image capture device as described above and can be or is implemented, for example and in particular, by a digital camera. Furthermore, it is possible to create a background model in another way. Therefore, the invention contemplates the artificial creation of background models. This can be achieved, for example and in particular, by computer-generating the background models by storing components of a background in a supply and assembling or combining them to form a background.
[0132] According to the invention, the components can be components from real environments or scenes, but also abstract shapes, textures and colored areas, points and lines and the like, which are combined to form a background model in order to generate a plurality of different background models in a computer-aided or computer-implemented or automated manner.
[0133] The invention also considers the option of providing the training models with abstract elements as additional components that interfere with object recognition in order to achieve improved recognition of object bodies by means of training on the basis of scene simulations with a rich variety of shapes and an increased degree of difficulty.
[0134] Therefore, the invention also makes it possible to combine abstract shapes or color areas with captured real scenes or fragments thereof in a background model.
[0135] For this purpose, a corresponding advantageous further development option of the invention provides that the background model is generated by providing a supply of shape or color surface fragments which are combined, in particular randomly or stochastically, to form a background model.
[0136] The invention thus achieves the advantage that any background constellation can be considered when training the AI module. The invention also allows for the option of creating the background using constellation rules that specifically control the construction of the background model.
[0137] This can be done, for example, using schemes or routines, such as a combination list or a script. Combination rules that specify constellations of, for example, and in particular, shapes, colors, textures, and the like, and their relative placement, can specifically control the creation of a pool of backgrounds. This can be advantageously used to train the AI module in a targeted manner and thus mitigate or eliminate possible or identified weaknesses in object recognition through adapted background models.
[0138] According to the invention, a background model can be created in a variety of ways. However, within the scope of the invention, it has proven advantageous to train the AI module on realistic constellations or scene arrangements.
[0139] Object recognition is particularly difficult in real-world scenes where a diverse array of elements and environmental influences occur. Therefore, it has been shown that realistic and varied training of an AI module leads to higher recognition rates and reliability.
[0140] Accordingly, it has been found within the scope of the invention that the use of an image of a real background as a background model favors the training of the Kl module.
[0141] The invention therefore considers that the background is digitally captured, or digitally recreated, or simulated, a real environment or scene. This allows images of real situations and environmental constellations to be used to train the AI module. Accordingly, corresponding influences from real environments or situations can be advantageously used when training the AI module to increase the reliability of object recognition according to the invention.
[0142] The invention takes into account that the two aforementioned options can be combined with each other and form a symbiosis.
[0143] Therefore, in an advantageous further development option of the invention, it is taken into account that a real background constellation is recorded in the background model.
[0144] The background constellation can also be achieved by combining fragments of captured real environments and situations.
[0145] According to the invention, the provision of a repository can be achieved, for example, and in particular, on the basis of databases and their connection for access, in particular bidirectional access, in which the object data sets and background data sets are stored in a retrievable manner. Within the scope of the invention, the individual object models or background models, as well as the scene arrangement, are provided with metadata, which in turn can be stored, for example, in a database, in order to make them available for further processing.
[0146] According to the invention, verifying the accuracy of the AI module can be made easier by including markers in the metadata that can be used as test criteria. This makes it possible to identify deviations in object recognition and identification and to obtain insights that can be used to create or generate scene arrangements for further training of the AI module.
[0147] This can be done based on customized routines that allow for adjustment based on the detected deviations. For this purpose, the invention takes into account that the AI module determines metadata that can be used to evaluate differences in object recognition and identification.
[0148] In order to verify the reliability of object recognition and the accuracy of identification in a cost-effective manner, a further advantageous development of the invention provides that the metadata include markers by means of which deviations in the recognition of object patterns or object models can be checked or are checked when verifying accuracy. Such markers can, for example, and in particular, be position markers by which the object models are positioned or aligned in a scene arrangement.
[0149] It has been shown within the scope of the invention that the detection of objects and the identification of object bodies are feasible on the basis of metadata. Therefore, the representations of a scene arrangement contained in the image data sets can be neglected when analyzing an image data set for a real scene and are ultimately dispensable. The relevant representations can be deleted after training the Kl module. Therefore, the invention also considers the option of deleting the image data sets of the scene arrangements within an iteration step. This reduces the memory requirement and leads to a reduction in system requirements. The invention can demonstrate its advantages, for example and in particular, when using a computer-implemented object detection method for user guidance. This can be the case for installation processes for electrical components and parts.Furthermore, the invention can be used for training and continuing education measures. Finally, the invention can advantageously support the automation of assembly processes as well as the monitoring and maintenance of electrical systems, for example and in particular control cabinets, by reliably detecting and identifying corresponding object bodies.
[0150] The invention is explained in more detail below with reference to the attached figures, in which, representative of a large number of variants of an inventive
[0151] An embodiment of the object recognition method is presented.
[0152] The figures are intended to illustrate the inventive
[0153] The corresponding steps or features can be derived from object recognition procedures.
[0154] All features claimed, described and shown in the drawings, taken individually and in any combination with one another, form the subject matter of the invention, regardless of their summary in the patent claims and their references, as well as regardless of their description or representation in the drawings.
[0155] Therefore, the features are not tied to the constellation explained below, but can also form an object recognition method according to the invention in isolation from one another as well as in a different combination or constellations.
[0156] The figures of the drawing show the aforementioned embodiment of the invention using schematic representations.
[0157] The representations in the figures are therefore not necessarily to scale, so that, among other things, the scales chosen in the figures may also differ from one another.
[0158] For clarity, the illustrations have been reduced to the elements that aid understanding. In the figures, identical or corresponding components / components or elements are provided with the same reference symbols.
[0159] For a better overview, not all elements / components / parts are always provided with reference symbols in the figures, whereby the assignment results from the same representation or a representation adapted to the view.
[0160] It shows:
[0161] Fig. 1 shows an embodiment of a computer-implemented object recognition method representative of a plurality of variants according to the invention which deviate therefrom,
[0162] Fig. 2 shows an example of a composition of a scene arrangement created according to the first embodiment of a computer-implemented object recognition method,
[0163] Fig. 3 shows a capture of a scene arrangement from different capture perspectives, in which the background model is projected onto a concave surface and object models are arranged opposite this surface,
[0164] Fig. 4 shows the use of markers for positions of object models based on the scene arrangement shown in Fig. 3,
[0165] Fig. 5 the scene arrangement shown in Fig. 3 with additional object models.
[0166] Fig. 1 shows detection steps 2 of a computer-implemented object recognition method 4 for identifying object bodies in a real scene. The detection steps 2 include
[0167] I. Providing an image data set of a
[0168] Real scene captured by the image capture device for evaluation by a Kl module,
[0169] II. Analysis of the image dataset by the Kl module and detection of object patterns in the real scene contained therein, III. Identification of object bodies contained in the real scene based on the detected object patterns by the Kl module,
[0170] In this embodiment, the detection steps 2 are set up and designed such that the identification takes place in real time.
[0171] The object recognition method 4 comprises a phase in which the Kl module is trained for this purpose through the following iteration steps 6: i. Providing a supply of object data sets for object models, ii. Providing a supply of background data sets for background models, iii. Providing a plurality of image data sets for a digital scene arrangement, iv. Analyzing the respective image data set and recognizing object patterns in the scene arrangement contained therein, v. Identifying the object model(s) contained in the scene arrangement, wherein a. at least one fragment of the object patterns and / or a background pattern is detected in the scene arrangement as a reference for determining a, b. the scale of the graphical representation is determined, c. the respective object pattern contained in the scene arrangement is recognized based on the scale, d. the respective object model is determined based on the respective object pattern, vi.Verifying the accuracy of identifying object bodies based on metadata, such as, for example, position, orientation, color, and light values, of the digital scene arrangement; vii. Repeating the iteration steps if the verification results in insufficient accuracy.
[0172] The iteration step iii includes further process steps and provides that the image data sets are created by means of a scene creation module through the following steps: a. Reading in a, in particular randomly selected,
[0173] background data set, b. reading in a, in particular randomly selected,
[0174] object data set or a plurality of, in particular randomly selected, object data sets, c. arranging the object model of the object data set or the object models of the object data sets and the background model of the background data set to form a scene arrangement, d. capturing the scene arrangement in a digital representation, e. storing the digital representation in an image data set and storing metadata for the digital representation of the scene arrangement,
[0175] A training of the Kl module according to the invention is typically carried out as a phase prior to the operational use of an object recognition method 4, wherein the iteration steps iii are not necessarily only carried out in an initiation phase for the object recognition method 4, but can also serve to carry out an extended training if deviations from the accuracy of the object recognition occur.
[0176] Therefore, the iteration steps can also be carried out at a later time after an initial training or can be carried out concurrently with or between the detection steps 2.
[0177] Furthermore, according to the invention, further steps can be linked to the process steps described within the scope of the invention, which are represented in Fig. 1 by a dashed box symbol. Fig. 1 schematically shows the detection steps 2 with which a scene arrangement 10 is created and recorded for training the AI module 12, as illustrated in Fig. 2.
[0178] The AI module 12 has a neural network 14 used for object recognition, which is trained using training data 16. The training data 16 includes previously captured scene arrangements 10 provided in image data sets 18 and the respective metadata 20, which characterize a scene arrangement 10, for example, in particular with regard to its components 22, such as the object models 24 contained therein with a background model 26, their arrangement and alignment relative to one another, and other parameters.
[0179] Further parameters are, for example and in particular, image and capture perspective settings with which the scene arrangement 10 is captured by means of a virtual image capture device or camera 28.
[0180] In this exemplary embodiment, the object models 24 required to construct a scene arrangement 10 are provided in respective object data sets 30, wherein the object model 24 contained therein is contained in a vector-based data format, specifically in a CAD data format in this exemplary embodiment.
[0181] To create the scene arrangement 10, a store 32 of object data sets 30 for object models 24 is first provided. The store 32 of object data sets 30 comprises a spectrum of differently designed object models 24 (each provided with the same reference numeral in Fig. 2). The object models represent electrical components, such as the terminal blocks shown in Fig. 2. Electrical components are diverse and available in various forms. In this exemplary embodiment, the store 32 of object data sets 30 is based on the actually available products of a company, which are digitally represented by the respective object models 24.
[0182] In this embodiment, the object models 24 comprise the spectrum of products of a product portfolio and originate from a product database 36. This makes separate storage for the object models 24 obsolete, since they are provided directly from the product database 36.
[0183] To create the scene arrangement, a pool 38 of background data sets 40 is provided for background models 26 (each provided with the same reference numeral in Fig. 2). The background data set 40 is configured and designed such that it effects exposure of the plurality of object models 24, 241, 242 in the scene arrangement 10.
[0184] The object models 24 as well as the background models 26 are each provided as parameterized models that can be adjusted to create the scene arrangement 10 and are recorded accordingly adjusted and arranged in a digital representation 42. In Fig. 2, a plurality of the object models 24 are arranged next to one another in a row arrangement 44, with one object model 24 serving as a base 46 for the row arrangement 44, on which the other object models 24 in the aforementioned row arrangement 44 are arranged.
[0185] Fig. 3 shows a schematic representation of a scene arrangement 10 in which object models are arranged opposite one another on a concave surface, and the background model is projected onto this concave surface. Fig. 3 illustrates that the iteration step of arranging c includes projecting the background model 26 contained in the randomly selected and read-in background data set 40 onto a curved surface 46, which in this embodiment is hemispherical.
[0186] For this purpose, the background model 26 contained in the background data set 40 is provided as a panoramic image, so that the aforementioned projection is simplified and improved.
[0187] The digital representation 42 is stored in the image data set 18 using an image data format that contains the background model 26 as a panoramic image.
[0188] In this embodiment, capturing the scene arrangement 10 includes a variation of a capturing perspective 48, wherein the variants of the capturing perspective 48 are captured in a respective digital representation 42.
[0189] The capture perspective 48 depends, among other things, on the setting of the virtual image capture device or camera 28 as well as on its position and orientation.
[0190] This is illustrated in Fig. 3 by the plurality of virtual detection devices or cameras 28.
[0191] The scene arrangement is captured in such a way that a variation of at least one of the following aspects: the background model 26, the selection of the object models 24, 241, 242, the arrangement of the object models 24, 241, 242 relative to one another, and the light or contrast values of the scene arrangement 10 is created, and the variants of the scene arrangement 10 are captured in a respective digital representation. Thus, a plurality of representations 42 or image data sets 18 are available for training the AI module.
[0192] According to the invention, it is possible to construct the background model from various components by creating a background model, for example and in particular by means of one of the following options:
[0193] • the background model 26 is generated by providing a supply of shape or color surface fragments, which are combined, in particular randomly or stochastically, to form a background model 26,
[0194] • a real background constellation is recorded in the background model 26.
[0195] The previously described embodiment does not make use of this option, although this can also be a design according to the invention.
[0196] The object recognition method 4 illustrated by the figures can also be designed in other forms according to the invention and in particular can use further features as contained in the patent claims or within the description and the figures.
[0197] For example, it is possible to create a scene arrangement 10 in which some components or parameters are kept constant and other components or parameters are varied for capture in different digital representations 42. This makes it possible, among other things, to train the Kl module to recognize size ratios or scales and to take these into account accordingly in object recognition.
[0198] Fig. 4 illustrates the possibility of detecting positions of markers and using these markers to test the accuracy of the Kl module 12 for object recognition.
[0199] Fig. 4 shows a schematic representation of the scene arrangement 10 from Fig. 2 without a representation of the background model 26.
[0200] Fig. 4 illustrates one possibility for markings 50 on object models 24 that can be used as test criteria. Using markings 50 according to the invention, it is possible, on the one hand, to represent a position of a respective object model 24, based on which the Kl module 12 can be checked for its accuracy in object recognition and identification of object bodies 24, 241, 242. For clarity, in Fig. 4, only one object model 24 is provided with the reference symbol 24 as a representative of the other models.
[0201] To this end, the respective markings 50 can be recorded in the metadata 20, by means of which deviations in the recognition of object patterns can be checked or checked when verifying the accuracy.
[0202] In this exemplary embodiment, the aforementioned markings 50 comprise bounding boxes 52 for an object model 26, which surround the object model 26 on at least three sides of the object model 26. In this exemplary embodiment, this is sufficient for examining the object pattern recognized in a real scene in order to detect a deviation that may be the cause of the recognition of object models 26 contained in the scene arrangement 10. In Fig. 4, the previously described marking is symbolized by a dashed frame and is identified by the reference numeral 52 as representative of further markings.
[0203] This ultimately makes it possible to adapt the training of the Kl module 12 according to the detected deviations. This can be done automatically according to the invention.
[0204] Fig. 5 illustrates, using the scene arrangement 10 shown in Fig. 3, a schematic representation of a problem that arises in the field of electrical engineering applications and scenarios. In Fig. 4, for clarity, only one object model 24 is designated with the reference numeral 24. Furthermore, in Fig. 5, the number of components of a scene arrangement is shown in reduced detail for clarity.
[0205] In the scene arrangement, which is modeled on a possible real scene, in addition to object models 24 arranged in a structured arrangement relative to one another, there are also further object models 241, 242 that represent electrical connectors for connecting the terminals 54 (uniformly designated by reference numeral 54) with equal electrical potential. Object model 242 represents a so-called jumper, and object model 241 represents a connecting cable, which, as a representative of flexible components of a real scene, can also be provided as a parameterized object model in order to adapt it to the respective scene arrangement.
[0206] This contributes to the fact that such components do not have to be stored in a plurality of object models 24 and, at the same time, a simple adaptation to various possible real situations is made possible according to the invention.
[0207] Many implementation variants of a computer-implemented object recognition method according to the invention are shown, so that the embodiment explained with reference to the figures is not limiting to the possible embodiments according to the invention.
[0208] List of reference symbols
[0209] Detection steps 1,11,111 Product database 36
[0210] Iteration steps i,ii,iii,iv,v,vi,vii, a,b,c,d,e Stock of background data sets 38
[0211] Detection steps 2 Background data set /
[0212] Object recognition methods 4 Background data sets 40
[0213] Iteration steps 6 Digital representation 42
[0214] Scene arrangement 10 Row arrangement 44
[0215] Kl-Module 12 Curved surface 46 Neural network 14 Detection perspective 48
[0216] Training data 16 Mark(s) 50
[0217] Image dataset / Image datasets 18 Bounding box 52
[0218] Metadata 20 Connections 54
[0219] Components 22
[0220] Object model(s) 24
[0221] Object models (connecting cables) 241
[0222] Object model (jumper) 242
[0223] Background model(s) 26
[0224] Virtual image capture device(s) /
[0225] Camera(s) 28
[0226] Object data set / object data sets 30
[0227] Stock of object records 32
Claims
Patent claims 1 . Computer-implemented object recognition method (4) for identifying object bodies in a real scene by the following detection steps (2): IV. Providing an image dataset of a image capture device captured real scene for evaluation by a Kl module (12), V. Analysis of the image data set with the captured real scene by the Kl module (12) and detection of object patterns in the real scene contained therein, VI. Identifying object bodies contained in the real scene based on the recognized object patterns by the Kl module (12), wherein the Kl module (12) is trained for this purpose by the following iteration steps (6): viii. Providing a store (32) of object data sets (30) for object models (24, 241, 242), ix. Providing a store (38) of background data sets (40) for background models (26), x. Providing a plurality of image data sets (18) for a digital scene arrangement (10), wherein the image data sets (18) are created by means of a scene creation module by the following steps: a. Reading in a, in particular randomly selected, background data set (40), b. Reading in a, in particular randomly selected, object data set (30) or a plurality of, in particular randomly selected, object data sets (30), c.Arranging the object model (24, 241, 242) of the object data set or the object models (24, 241, 242) of the object data sets (30) and the background model (26) of the background data set (40) to form a scene arrangement (10), d. Capturing the scene arrangement (10) in a digital representation (42),. e. Storing the digital representation (42) in an image data set (18) and storing metadata (20) for the digital representation (42) of the scene arrangement (10), xi. Analyzing the respective image data set (18) and recognizing object patterns in the scene arrangement (10) contained therein, xii. Identifying object models (24) contained in the scene arrangement (10), wherein a. at least one fragment of the object patterns and / or a background pattern is detected in the scene arrangement as a reference for determining a scale, b. the scale of the graphical representation is specified, c. the respective object pattern contained in the scene arrangement is recognized based on the scale, d. the respective object model is determined based on the respective object pattern, xiii. Verifying an accuracy for identifying object bodies based on the metadata (20) for the digital scene arrangement (10), xiv.Repeat the iteration steps if the verification results in insufficient accuracy.
2. Computer-implemented object recognition method (4) according to claim 1, characterized in that the object data set (30) has a vector-based data format, in particular a CAD data format.
3. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the stock of object data records comprises a spectrum of differently designed object models, in particular the spectrum of the products of a product portfolio, preferably the object data records originate from a product database.
4. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the object data set (30) contains the object model (24, 241, 242) in an outer surface and / or outer edge-reduced representation.
5. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the parameters for detecting a scene arrangement (10) are varied before the digital representation (42) is stored in an image data set (18).
6. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the object models (24) or the background models (26) are provided as parameterized models which can be adjusted for the creation of a scene arrangement (10).
7. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the arrangement of a plurality of object models (24, 241, 242) to one another comprises at least one of the following forms: • Row arrangement, • Column arrangement, • Matrix arrangement, • equally spaced arrangement.
8. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the arranging comprises projecting the background model (26) contained in the background data set (40) onto a curved, in particular concavely curved, preferably a hemispherical surface (46).
9. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the background model (26) contained in the background data set (40) is provided in the manner of a panoramic image.
10. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the Background data set (40) is set up and designed such that it effects an exposure of the object model (24, 241, 242) or the plurality of object models (24, 241, 242) in the scene arrangement (10). 1 1. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that at least the digital representation (42) is stored in an image data set (18) with an image data format, wherein this image data set (18) contains in particular the background model (26) as an image, in particular as a panoramic image, preferably as a 360-degree panoramic image.
12. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the detection of the scene arrangement (10) comprises a variation of the detection perspective (48), wherein the variants of the detection perspectives (48) are detected in a respective digital representation (42).
13. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the detection of the scene arrangement (10) is carried out in such a way that a variation of at least one of the following aspects • the background model (26), • the selection of object models (24, 241 ,242), • the arrangement of the object models (24, 241, 242) to each other, • the light or contrast values of the scene arrangement (10) are created and the variants of the scene arrangement (10) are recorded in a respective digital representation (42).
14. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the background model (26) is created by means of at least one of the following options: • the background model (26) is generated by providing a supply of form or color surface fragments, which, in particular randomly or stochastically, to form a background model (26), • a real background constellation is recorded in the background model (26).
15. Computer-implemented object recognition method (4) according to one of the preceding claims, characterized in that the metadata (20) comprise markings (50) by means of which deviations in the recognition of object models (24) can be checked.
Citation Information
Patent Citations
Techniques for training machine learning
US20200302241A1
Generating Synthetic Image Data for Machine Learning
US20200342652A1
Synthetic dataset creation for object detection and classification with deep learning
WO2023014369A1