Device designed for acquiring images and / or videos of points and areas of support of a subject on a surface

A device with a transparent support and camera, combined with a convolutional neural network, addresses the need for accurate image and video capture of baby support points, enabling objective assessment and timely intervention for motor development delays.

WO2026099527A1PCT designated stage Publication Date: 2026-05-15BIOMEDICAL RES INST OF SALAMANCA OF THE HEALTH SCI INST OF CASTILLA Y LEÓN +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BIOMEDICAL RES INST OF SALAMANCA OF THE HEALTH SCI INST OF CASTILLA Y LEÓN
Filing Date
2025-11-04
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

There is a need for devices that can accurately acquire images and/or videos of a baby's support points and areas on a surface to assess posture, which is crucial for early detection of motor development delays to facilitate timely interventions.

Method used

A device comprising a rectangular frame with a transparent support and a camera positioned at the intersection of the frame and ground, capable of capturing high-resolution images and videos under ambient lighting, combined with a convolutional neural network model for image segmentation into multiple classes, allowing for objective and quantitative assessment of motor development.

Benefits of technology

Enables robust detection of support zones with high accuracy, facilitating timely and precise interventions by identifying potential developmental issues through continuous video capture and automatic image segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure ES2025070668_15052026_PF_FP_ABST
    Figure ES2025070668_15052026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention pertains to the field of medicine and relates specifically to a device designed for acquiring images and / or videos of points and areas of support of a subject on a surface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] DEVICE CONFIGURED FOR THE ACQUISITION OF IMAGES AND / OR VIDEOS OF THE SUPPORT POINTS AND AREAS OF A SUBJECT IN

[0003] A SURFACE

[0004] FIELD OF INVENTION

[0005] The present invention pertains to the medical field. Specifically, the present invention relates to a device configured for acquiring images and / or videos of the points and areas of contact of a subject on a surface.

[0006] STATE OF THE ART

[0007] Given the importance of postural control, the assessment of posture in babies becomes a crucial aspect within the evaluation of motor development.

[0008] In this assessment, healthcare professionals trained in pediatrics can identify early on possible delays or abnormalities in motor development that could limit the child's ability to explore their environment effectively, efficiently, and early.

[0009] This, in turn, could have implications for other areas of development, such as cognitive and social development. Early and accurate detection of posture allows for timely intervention with early interventions that can improve not only the baby's motor development but also their overall interaction with the environment. This is crucial to ensuring that all children have the best opportunities to reach their full developmental potential.

[0010] Therefore, there is a need in the state of the art for devices that can help acquire images and / or videos of the baby's support points and areas on a surface and that allow the baby's posture to be assessed.

[0011] The present invention focuses on solving this problem by designing a device configured for the acquisition of images and / or videos on the support points and areas of a subject on a surface.

[0012] BRIEF DESCRIPTION OF THE INVENTION

[0013] The present invention relates to a device configured for acquiring images and / or videos of a subject's points and areas of support on a surface. This invention has significant clinical implications, as it allows for an objective and quantitative assessment of the subject's motor development, which can help identify potential early developmental problems and facilitate more timely and precise interventions.

[0014] The first aspect of the present invention relates to a device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, comprising: a structure (2) formed by a rectangular frame and support points of the rectangular frame with the ground, a transparent support (3) supported on the rectangular frame on which the subject would be placed, characterized in that the device (1) comprises a support (5) for a camera (6) located at the intersection of the support points of the rectangular frame with the ground and under the transparent support (3).

[0015] In a preferred aspect, the device (1) comprises stops (4) through which the transparent support (3) rests on the rectangular frame.

[0016] In a preferred aspect, the device (1) is characterized in that each support point of the rectangular frame with the ground forms a right angle, so that it has a section that extends horizontally along the ground and another section that rises vertically to meet the rectangular frame.

[0017] In a preferred aspect, the device (1) is characterized in that it comprises two support points of the rectangular frame with the ground, each of them U-shaped, or four support points of the rectangular frame with the ground, each of them L-shaped.

[0018] In a preferred aspect, the device (1) is characterized in that it comprises a camera and a camera support that is located at the intersection of the support points of the rectangular frame with the ground and under the transparent support, such that it is located in the center of the section that extends horizontally along the ground.

[0019] In a preferred aspect, the device (1) is characterized in that the rectangular frame is square.

[0020] In a preferred aspect, the device (1) is characterized in that the subject is a person, preferably a baby or a child.

[0021] In a preferred embodiment, the device (1) is characterized in that the subject is positioned on the transparent support (3) in a supine and prone position. In a preferred embodiment, the device (1) is characterized in that the transparent support (3) has corner protectors (7).

[0022] In a preferred embodiment, the device (1) is characterized in that the transparent support (3) is made of glass or an optically transparent equivalent. The present invention expressly describes a glass platform used as a support surface during image capture. This demonstrates that the material must possess optical properties that allow the observation and recording of the subject's contact areas from beneath the support. Therefore, the claim protects not only the glass itself, but also any glass or optically transparent equivalent, such as tempered glass, methacrylate, or polycarbonate, that provides the same technical effect of transparency and rigidity necessary for obtaining high-quality images.

[0023] In a preferred embodiment, the device (1) is characterized in that it is configured for the continuous capture of high-resolution video or successive images. The present invention details that the acquisition of visual information is performed by means of continuous video recording, from which frames are subsequently extracted to compose the training and validation dataset. This configuration ensures the acquisition of a sufficiently dense temporal sequence of images to analyze the subject's movement and support in real time.

[0024] In a preferred aspect, the device (1) is characterized in that the acquired images have sufficient resolution to identify the subject's contact areas and are stored in a standard digital format. The experimental example includes the generation of 2704x2028 pixel RGB color images and their storage in JPEG format. These values ​​demonstrate that a resolution of this order is sufficient to accurately distinguish the contact areas between the subject and the transparent support.

[0025] In a preferred embodiment, the device (1) is characterized in that image acquisition is performed under ambient lighting conditions without requiring controlled and structured light sources. The present invention expressly states that the capture was carried out under room lighting conditions, without resorting to controlled and structured light. This configuration simplifies the system and its installation and demonstrates that the device can operate effectively with ordinary ambient light, maintaining good visibility of the support areas.

[0026] In a preferred embodiment, the device (1) is characterized in that the subject is an infant or young child positioned on the transparent support (3). The present invention describes the use of the device with infants from 0 to 6 months of age placed in prone and supine positions on the glass support. Comparative analyses between different ages (under 3 months and around 6 months) are included, demonstrating the applicability of the device to infants and young children, the population in which postural development is to be studied.

[0027] In a preferred embodiment, the device (1) is characterized by its robustness in detecting support points against variations in lighting conditions, the presence of accessories, or anthropometric differences in the subject. The present invention specifies that the results were consistent regardless of diaper use, variations in weight or height, the number of people in the room, or lighting conditions. This behavior confirms that the device and the capture method offer robustness against external disturbances, ensuring the reliability of support detection in real-world environments.

[0028] The present invention also relates to a system for automatically identifying the support zones of a subject placed on a transparent support (3), comprising the device according to any of the preceding claims and a processor configured to perform image segmentation into multiple classes or regions of interest. The present invention describes a system that combines the physical device with digital image processing using multi-class segmentation techniques, specifically into six classes (ceiling, frame, glass, person, supported, and unsupported).

[0029] In a preferred embodiment, the system is characterized in that the processor uses a convolutional neural network model trained by transfer learning. In the described embodiment, a ResNet-50 pre-trained in ImageNet is employed, adapted by transfer learning to the specific problem classes and the generated image dataset. This example demonstrates the application of a convolutional neural network (CNN) model capable of generalizing to new subjects or conditions, which constitutes the protected principle of this claim. In a preferred embodiment, the system is characterized in that the trained model achieves a level of accuracy sufficient to distinguish the subject's support zones with high reliability.The experimental results of the document show recall values ​​greater than 90%, loU greater than 0.76 and MeanBFScore greater than 0.80 in all classes, demonstrating the reliability of the model for segmenting support zones.

[0030] In a preferred embodiment, the system is characterized in that the model is configured to distinguish at least between contact regions and non-contact regions of the subject with the transparent support. The present invention includes among its classes the categories Supported and Not Supported, and the confusion matrix shows that 91.8% of the pixels predicted as Supported are correct, while only 7% belong to Not Supported. These figures illustrate the system's ability to effectively distinguish contact and non-contact areas, an objective that defines the functional essence of this claim.

[0031] The present invention also relates to a method for generating a dataset for training and validating a machine learning model for segmenting images of subjects on a transparent surface, comprising: (i) capturing video sequences; (ii) extracting frames or images of the subject; (iii) labeling the images by identifying regions corresponding to different classes; and (iv) using these images to train, validate, and test the model. The present invention describes step by step the creation of the dataset, the capture of a free-moving video of the baby, the extraction of frames, the annotation of six classes in each image, and the subsequent division of the dataset into 60 / 20 / 20% for training, validation, and testing. This sequence exemplifies the general protected method, which is applicable to any type of segmentation model.

[0032] In a preferred embodiment, the procedure is characterized in that the annotation of the images is performed manually or semi-automatically by expert personnel. The present invention details that the manual annotation of the Ground Truth was performed by physiotherapists, coloring each pixel of the images to define the six classes.

[0033] In a preferred embodiment, the procedure is characterized by the fact that the model is trained until convergence or stability of the learning parameters is achieved. During the documented training, the accuracy and loss curves show convergence around iteration 1,200, demonstrating that the criterion for terminating the learning is based on parameter stability.

[0034] The present invention also relates to the use of this system to identify the support areas of a subject in different positions on the transparent support. The present invention provides examples in prone and supine positions, analyzing the body's support points (trunk, head, pelvis, limbs). These examples demonstrate that the system can be used to identify support areas in different postures.

[0035] The present invention also relates to the use of this system to compare automatically identified support zones with those assessed by a qualified professional. The comparison between the network's predictions and the opinions of physiotherapists is explicitly described in the present invention, where the correlations between the automatic result and the expert's manual assessment are analyzed. This application allows for the clinical validation of the system.

[0036] The present invention also relates to a computer program containing instructions executable by a processor for segmenting images or videos according to the procedure described above. The document describes the computer implementation of the trained model, stored in an executable file, which allows for the automatic segmentation of new images into the six established classes, with an average processing time of approximately 6 seconds per image. This specific example supports the protected concept of an executable computer program for applying the described segmentation procedure.

[0037] For the purposes of the present invention, the following shall be understood:

[0038] • “Transparent support” (3): An optically transmissive structural element, preferably made of vitreous material, that allows observation of the contact surface from its underside. This includes glass and other equivalent optically transparent materials, such as tempered glass, methacrylate, or polycarbonate. “Subject” (B): A person placed on the transparent support, preferably an infant or young child, although the term includes any human being or organism whose support pattern is to be recorded or analyzed.

[0039] • “Support zone or area”: region of the subject’s body that maintains physical contact with the transparent support, visually identifiable in the image or video by means of optical contrast or automatic segmentation.

[0040] • “Continuous capture of video or successive images”: process of acquiring visual data in a temporally sequenced manner, either by video or by taking multiple consecutive high-resolution images.

[0041] • “Image segmentation into multiple classes or regions of interest”: a digital processing operation that assigns to each pixel of the image a label belonging to one or more categories (e.g., glass, person, supported, not supported).

[0042] • “Convolutional neural network model trained by transfer learning”: a training technique or strategy within artificial intelligence that reuses the weights and structures of a previously trained network (e.g., ResNet-50 on ImageNet) and adjusts them to the specific dataset generated with the device.

[0043] • “High reliability”: a model is considered to achieve high reliability when it consistently and accurately reproduces the identification of support zones observed by expert personnel, as evidenced in the documented examples with accuracy values ​​above 90%.

[0044] • “Convergent or stable learning”: the point in training where the loss or error metrics stop varying significantly, exemplified in the document around 1,200 iterations.

[0045] • “Expert or qualified personnel”: professionals with technical or clinical knowledge in postural observation (e.g., physiotherapists) responsible for performing the annotation or manual validation of the images.

[0046] • “Contact / non-contact regions”: areas of the body classified respectively as Supported and Unsupported, detected through visual contrast or model inference. Description of the figures

[0047] The device of the invention is illustrated more specifically in the Figures, where the following technical characteristics have been reflected according to the following numbering:

[0048] Figure 1. Top view of device (1) comprising a structure (2) formed by a rectangular frame and support points of the rectangular frame with the ground, a transparent support (3) supported on the rectangular frame by means of stops (4) on which the subject would be placed, a support (5) for a camera (6) located at the intersection of the support points of the rectangular frame with the ground and below the transparent support (3).

[0049] Figure 2. Side view of a device comprising a structure (2) formed by a rectangular frame and support points of the rectangular frame with the ground, a transparent support (3) supported on the rectangular frame by means of stops (4) on which the subject (B) would be placed, a support (5) for a camera (6) located at the intersection of the support points of the rectangular frame with the ground and below the transparent support (3) and protectors (7) of the corners of the transparent support (3).

[0050] Figure 3. Images taken from videos of three sessions showing babies in prone position with two different physiotherapists and other people. Figure 4. Frame painted by an expert.

[0051] Figure 5. Comparison of frames segmented by experts (Ground Truth) and by CNN (Baby 3).

[0052] Figure 6. Comparison of frames segmented by Ground Truth experts and by CNN (Baby 26).

[0053] Figure 7. Comparison of real and segmented frames by CNN and expert criteria (Baby 3). Physiotherapist comments: 102. AC: not detected because it is elevated. AT: accurately detects support. AA: does not detect full abdominal support, only upper support. AP: does not detect support due to the diaper. AMMSS: accurately detects support. AMMII: does not detect support due to the baby's diaper. This area is detected less accurately. 302. AC: not detected because it is elevated. AT: accurately detects support. AA: does not detect full abdominal support, only upper support. AP: does not detect support due to the diaper. AMMSS: detects very little support area. AMMII: does not detect support due to the baby's diaper. This area is detected less accurately. 402. AC: not detected because it is elevated. AT: accurately detects support. AA: does not detect full abdominal support, only upper support. AP: does not detect support due to diaper.AMMSS: only detects the support area on the right hand. AMMII: does not detect support due to the baby's diaper. This area is detected with less accuracy. 752. AC: detects a very small support area and does not detect the support of the evaluator's hands. AE: does not detect full support of the lower lumbar spine, only of the thoracic and upper lumbar spine. AG: does not detect support due to the diaper. AMMSS: does not detect support of the elbows. AMMII: does not detect support because it does not exist. 852. AC: detects a very small support area. AE: does not detect full support of the lower lumbar spine, only of the thoracic and upper lumbar spine. AG: does not detect support due to the diaper. AMMSS: minimally detects support of the right arm. AMMII: does not detect support because it does not exist. 1102. AC: detects a very small support area and does not detect support of the evaluator's hands. AE: does not detect full support of the lower lumbar spine, only of the thoracic and upper lumbar spine.AG: does not detect support due to the diaper. AMMSS: does not detect arm support. AMMII: does not detect support because it does not exist.

[0054] Figure 8. Comparison of real and segmented frames by CNN and expert criteria (Baby 26). Physiotherapist comments: 102. AC: does not detect head support because the baby has not yet been placed on the platform. AE: only detects support of the lower lumbar region because the rest has not yet been placed on the platform. AG: accurately detects support of both buttocks. AMMSS: does not detect arm support because it is not present; the baby is suspended. AMMII: only detects support of the left thigh; the rest cannot be detected because the baby is suspended. 302. AC: correctly detects the baby's head support. AE: accurately detects support of the baby's entire back. AG: accurately detects support of both buttocks. AMMSS: detects support of the baby's right arm; the left arm cannot be detected because it is raised. AMMII: accurately detects thigh support. 402.AC: Correctly detects the baby's head support. AE: Accurately detects the support of the baby's entire back. AG: Accurately detects the support of both buttocks. AMMSS: Detects the support of the baby's right arm; the left arm cannot be detected because it is raised. AMMII: Accurately detects the support of the right thigh and part of the left thigh that is supported. 777. AC: Does not detect the baby's face because it is raised. AT: Does not detect the support of the trunk because it is suspended in the evaluator's hands. AA: Detects the support of the abdomen of the part positioned on the platform with considerable accuracy. AP: Detects support with considerable accuracy. AMMSS: Detects the support of the baby's right hand and part of the fingers of the left hand. AMMII: Correctly detects the support of the medial aspect of both thighs. 852. AC: Does not detect the baby's face because it is raised. AT: Does not detect the support of the trunk because it is raised.AA: Detects abdominal support quite accurately. AP: Detects support quite accurately. AMMSS: Detects support of the palm and wrist of the right hand, as well as the baby's left arm and forearm. AMMII: Correctly detects support of the medial aspect of both thighs, especially the left thigh, which is more firmly supported on the platform. 1102. AC: Does not detect the baby's face because it is elevated. AT: Detects support correctly. AA: Detects abdominal support quite accurately. AP: Detects support quite accurately. AMMSS: Accurately detects support of both forearms. AMMII: Correctly detects support of the medial aspect of both legs.

[0055] Figure 9. Normalized confusion matrix.

[0056] Figure 10. Results of the comparison of image 35.

[0057] Figure 11. Comparison of the caudal support area in two infants. Left: The green support area is distributed throughout the body. The support and righting surface is not competent; there are no support or righting points because the infant is less than 3 months old and at this age is not yet able to establish support points that would allow them to shift their center of gravity caudally. Right: The green support area shifts caudally, to the level of the pubic symphysis and lower limbs, because the infant is able to right themselves on their elbow and forearm, respectively, and raise their head against gravity, shifting their weight and center of gravity caudally. This establishes competent support and righting points, allowing them to raise their head against gravity and establish adequate head control.

[0058] Detailed description of the invention

[0059] The present invention is illustrated by means of the following Examples which explain the application of the device of the invention.

[0060] Example 1. RESEARCH METHODOLOGY

[0061] Example 1.1. DESIGN

[0062] This is an observational and cross-sectional study.

[0063] Example 1.2. SAMPLE / PARTICIPANTS

[0064] The sample consists of infants aged 0 to 6 months. A non-probability convenience sampling method was used.

[0065] These infants were evaluated at a Pediatric Physiotherapy Center in Salamanca. The study has been approved by the Ethics Committee of the University of Salamanca under registration number 840.

[0066] The study's objectives and procedures were disseminated through health centers, hospitals, and social media, among other channels, using an informational poster with contact information to recruit families interested in participating. Inclusion criteria: Healthy infants aged 0 to 6 months. Infants whose parents signed a consent form to participate in the study and authorized the taking of images and video.

[0067] • Exclusion criteria: o Babies under 37 weeks old. o Babies under 6 months old with a described pathology or diagnosed motor delay.

[0068] It is important to note that this study did not analyze the posture of premature babies, whose motor development is deprived of uterine containment for a different period of time depending on the degree of prematurity.

[0069] Example 1.3. VARIABLES

[0070] The main variable of study will be posture, which will be evaluated in dorsal and ventral decubitus positions using the following instruments.

[0071] Example 1.3.1. Measuring instruments

[0072] Medical record

[0073] Sociodemographic data about the mother, the pregnancy, and the delivery were collected on a medical record sheet. Sociodemographic data about the baby, its current situation, and its sleep and feeding habits were also recorded.

[0074] Device of the invention

[0075] A prototype was designed to assess infant posture in supine and prone positions. This device is designed to hold infants while capturing images and videos of their support points and areas of contact in both positions. It consists of a transparent support and a camera, positioned at the bottom, for video recording. This device is used to analyze the infants' support areas. The support must meet several design requirements, including: the safety and comfort of the infants, an appropriate ergonomic position for professionals to perform the assessment, and a defined outline made of a non-reflective material.

[0076] Deep Learning Software Environment

[0077] A platform incorporating Deep Learning (DL or Deep Learning) libraries has been used to develop all the learning tasks in the DL workflow, such as:

[0078] • Construct the dataset, preprocess the input data to train and validate the convolutional neural network (CNN), since no dataset exists to our knowledge. This requires decomposing the videos into frames, which will later be processed to generate the Ground Truth (using the application described in the following section).

[0079] • Use pre-trained CNNs (in the context that is usually called transfer learning or building a new CNN).

[0080] • Train the network with configurable options and training loops, to adjust the model parameters to the specific problem.

[0081] • Validate the network's behavior during training.

[0082] • Subsequently evaluate the trained model using images different from those used in training.

[0083] Once training and validation are complete, the CNN model is deployed. This process involves exporting the trained and validated model for use in a specific application, via a processing device such as a GPU (Graphics Processing Unit) or a CPU (Central Processing Unit). In this way, the model can be integrated with other components to build a DL-based system for, for example, real-time image segmentation.

[0084] Image editing application

[0085] A software application for creating and editing digital images and artwork was used, available for iOS and macOS devices. This application features an intuitive interface and an image editing tool that allows working with layers, which is essential for creating and manipulating complex images. Furthermore, it offers a suite of advanced editing tools and supports a wide variety of formats, including JPEG, which was used in this project.

[0086] Example 1.4. WORK METHOD

[0087] A rehabilitation doctor, four physiotherapy researchers, and two Artificial Intelligence researchers participated in defining and designing the work method.

[0088] The method consists of several stages:

[0089] • Sample recruitment: This process began in January 2019. The study objectives and procedures were disseminated through an informational poster with contact information to recruit families interested in participating. The poster was distributed throughout the Salamanca University Hospital Complex (in the Pediatrics and Obstetrics wards), as well as in Salamanca Health Centers. It was also distributed to childbirth preparation groups, pediatricians, nurses, physiotherapists, and occupational therapists in the city of Salamanca, and shared via social media platforms such as WhatsApp, Facebook, and Instagram. The aim was to reach the largest possible target population to form the analysis sample.

[0090] Families interested in having their babies participate in the study called and / or contacted us via the email address included on the poster to schedule an appointment.

[0091] • Appointment for kinesiological assessment, medical history, and video recording using the device of the invention at the Pediatric Physiotherapy Center: The baby's family was invited to the Pediatric Physiotherapy Center to meet with the rehabilitation physician, the physiotherapy team, and the IT team. On the scheduled day, each family was informed personally and in detail about the research study. If the family consented before the assessment and video recording, they were given the following forms to complete: o The research study information sheet and Personal Data Protection form. o The authorization form for the minor's family to record images and videos for the study. • Recording of the baby's medical history: This included data relevant to the postural assessment.

[0092] • Kinesiological assessment procedure of the postural ontogenesis (motor and postural development) of the baby: it was carried out by the rehabilitation doctor and completed by the physiotherapy researchers.

[0093] • Video recording of the baby in prone and supine positions in the device of the invention.

[0094] • Building the dataset to subsequently train the CNN: the processing of the frames (a specific image within a sequence of moving images) is carried out with the image editing application to define the different characteristic facts.

[0095] • Design, training and deployment of the CNN, by researchers in Artificial Intelligence.

[0096] • Construction of the posture analysis support system based on deep learning techniques by integrating the CNN with the device of the invention.

[0097] • Analysis of the frames.

[0098] • Collection and management of results obtained after the analysis.

[0099] • Statistical analysis of the results obtained and the performance of the CNN.

[0100] Example 1.4.1. Kinesiological assessment

[0101] Once the medical history form was completed, the rehabilitation physician performed the kinesiological assessment of the babies. The baby was undressed, placed on a treatment table, and while the assessment was being carried out, the relevant data was collected on the assessment form designed to facilitate data collection by the physiotherapists.

[0102] The kinesiological assessment of motor development was based on the neurokinesiological diagnosis of Professor Václav Vojta, in which postural and motor development are evaluated as part of a global pattern within postural ontogenesis, and based on the level of uprighting achieved at each motor milestone. In this systematic approach, the baby's spontaneous motor skills were assessed first in prone and supine positions.Subsequently, each of the seven postural reactions, which quantify the level of righting (traction reaction, Landau reaction, axillary suspension reaction, Vojta lateral loss of balance reaction, Collis horizontal lateral suspension reaction, Peipert and Isbert vertical suspension reaction, and Collis vertical suspension reaction), was assessed to determine the quality of the infant's overall posture and thus relate it to the "ideal pattern," representing the maximum achievable postural quality. Primitive reflexes were also evaluated, relating them to the infant's chronological age by considering the persistence time of each primitive reflex (orofacial, tonic, cutaneous, and osteotendinous reflexes, among others).

[0103] The systematic approach to kinesiological assessment is based on the analysis of postural ontogenesis, which encompasses the innate postural and motor patterns that develop during the first year of a baby's life. The study of ontogenesis differentiates between patterns that develop from the prone position and those that develop from the supine position. As previously mentioned, the prone position is analyzed to evaluate the motor patterns involved in righting and gravity support functions; and the supine position is studied to evaluate the motor patterns related to grasping.

[0104] Example 1.4.2. Image capture in the device of the invention

[0105] Subsequently, video recordings were made of the infants in prone and supine positions using the device of the invention. When positioning the device, the lighting of the scene was carefully considered to avoid shadows, glare, or reflections, as excessive exposure could negatively impact image quality and the ability to extract features of interest. The recording camera was a remote-controlled video camera with features such as high definition, wide-angle lens, and remote control via a mobile device or computer. The rehabilitation physician or one of the physiotherapy researchers positioned the infant in both prone and supine positions.

[0106] It is essential to emphasize that, during the time the baby remained in the device of the invention, no type of stimulus or toy was offered to direct or attract its attention. Instead, the baby was allowed to move freely so that it could spontaneously regulate its posture on top of the device of the invention without receiving visual, auditory, perceptual, or tactile stimuli. Example 1.4.3. Construction of the dataset for CNN training

[0107] Once the video was captured, the frames were automatically extracted.

[0108] The objective of this phase was to obtain a complete dataset that allows the CNN to be trained. Each element of the training dataset consists of an original frame of the baby and the segmented image, where the characteristic features of the scene are colored / differentiated: glass, ceiling, frame, people, and other objects that the experts consider relevant.

[0109] This information is called Ground Truth in the field of CNN research. It is one of the most laborious and time-consuming tasks in a DL project, as it involves manually labeling, using appropriate software applications, each object with semantic meaning in each scene. Defining Ground Truth is a vital step because systematic inaccuracies in the training data will translate into inaccuracies in the CNN results.

[0110] In this method, the following characteristic facts to be identified have been defined:

[0111] 1. Area external to the device of the invention without significant facts.

[0112] 2. Metallic surface of the device of the invention.

[0113] 3. Internal area of ​​the device of the invention corresponding to the glass.

[0114] 4. The evaluators and the baby's parents, if they appear in the image.

[0115] 5. The baby's body.

[0116] 6. The support that the baby provides on the glass surface.

[0117] The researchers then colored these characteristic features using image editing software on each frame. This transformed the original image into the corresponding interpreted (segmented) image from the dataset (Ground Truth). For each frame, seven layers were created using the following color code.

[0118] 1. Layer 1 contains the original image as a base for coloring the other six layers.

[0119] 2. Layer 2 contains the entire surface of the image in gray. 3. Layer 3 defines the outer contour of the device of the invention and will be filled in blue.

[0120] 4. In layer 4, the internal contour of the device of the invention is delimited and will be filled in red.

[0121] 5. Layer 5 covers the entire surface of the evaluators and the baby's parents; if they appear in the image, it will be marked in orange.

[0122] 6. Layer 6 comprises the surface of the baby's body that was marked in yellow.

[0123] 7. Finally, in layer 7, the surface of the support that the baby makes on the device of the invention was marked in green.

[0124] Once all the layers were superimposed in the specified order, the resulting image contained only the significant facts. This final image was exported in JPG format.

[0125] For each video used in training, manual image editing was required, involving the automatic extraction of individual frames. The result of this phase is a set of (Real Image / Segmented Image) pairs that form the necessary dataset for training the CNN. Datasets typically contain several hundred or even thousands of images with sufficient resolution to ensure training quality.

[0126] Example 1.4.4. CNN Design, Training and Deployment

[0127] The typical workflow tasks in DL were performed, such as selecting the pre-trained network, training and validating it with the previously constructed dataset, and the final deployment of the model. As mentioned, the CNN was trained and deployed using the software platform that integrates DL libraries, and a GPU was used to significantly reduce training time.

[0128] CNNs are made up of multiple layers of artificial neurons. The advantage of a deep neural network is its ability to automatically learn meaningful low-level features, such as lines or edges, and merge them with higher-level features, such as shapes, in subsequent layers. A key property of images is taken into account: pixels that are close together are more strongly correlated than pixels that are farther apart. Mathematical convolution allows these relationships to be discovered, hence the name: Convolutional Neural Networks. In this way, CNNs exploit this property by extracting local features that depend only on small subregions of the image.

[0129] The information from these features can be fused in later processing stages to detect higher-order features and ultimately obtain information about the entire image. This is the key element of semantic segmentation. Furthermore, local features that are useful in one region of the image are likely to be useful in other regions, such as when the object of interest moves. By maintaining these local spatial relationships, CNNs are well-suited for image recognition tasks.

[0130] CNNs have been applied to a wide variety of tasks, including image classification, object localization and detection, image segmentation, and image registration. These networks allow for the capture of important feature relationships in an image (such as how pixels at an edge join to form a line) and reduce the number of parameters the algorithm has to compute, thus increasing computational efficiency. CNNs can take as input and process both two-dimensional and three-dimensional images with minor modifications to their architecture. In the field of health informatics, data analysis has experienced rapid growth due to the vast amount of multimodal data available. This has fueled interest in machine learning, especially deep learning, a technique based on artificial neural networks.

[0131] In this phase, as established in supervised learning, the calculations for training the CNN were performed using configurable options and training loops, which allowed the model parameters to be adapted to the specific problem. This is an iterative process where a trial-and-error strategy with information guides decision-making in different aspects such as network architecture, learning parameters, and the maximum resolution at which it can operate.

[0132] To reduce training times in a computationally intensive case like this, a specific software platform with enormous computational and visualization resources was used. Specifically, this software platform integrates the CUDA and cuDNN libraries that have enabled the development and viability of CNN technology by allowing it to leverage the power of NVIDIA GPUs and their parallel processing capabilities. This is essential for training CNNs, as well as for their execution under real-time conditions. It should be noted that learning experiments are tests that can last several days, even on high-performance computers.

[0133] Once the CNN was built and trained, its generalization ability was evaluated. This was done using a test set (TestSef) containing images not used during training, to check whether the network's predicted results were inaccurate.

[0134] Example 1.4.5. Construction of the posture analysis support system based on deep learning techniques

[0135] In this stage, the integration of the trained CNN with the device of the invention was addressed to build the posture analysis aid system based on deep learning techniques, which consists of modules or parts.

[0136] The initial part corresponds to the device of the invention, which is responsible for capturing the frames when the rehabilitation doctor or the physiotherapist places a new subject on the device of the invention in ventral and dorsal decubitus positions.

[0137] In the central part is the CNN to perform the segmentation of these images, automatically identifying the characteristic facts for which it has been trained, mainly the baby's supports on the device of the invention.

[0138] In this method for assisting in the analysis of the posture of subjects based on deep learning techniques, the real image of the subject in a frame is automatically transformed into a segmented image, where the different defined classes appear colored and, especially, the supports of the subject on the object of invention.

[0139] The final section displays the results. On the monitor, the physiotherapist can compare the actual image with the colorized image, where they can see the supports highlighted in green, which helps them identify possible deviations and irregularities in the infant's structures.

[0140] Example 2. DEVICE VALIDATION

[0141] The purpose of this example is to validate the technical operation of the described device and segmentation system, demonstrating that they allow for the objective and reproducible identification of support areas.

[0142] Example 2.1. Creating the dataset

[0143] The dataset will enable the training of CNNs (Convolutional Neural Networks), a type of deep, multi-layered neural network specialized in image segmentation, with the aim of identifying the support zones of infants aged 0 to 6 months on a glass platform. This involves the construction of a dataset specifically for this application domain. No other dataset with the characteristics described below is known to exist.

[0144] The data source for the dataset is the video captured while a physiotherapist positions and holds the baby on the platform. The baby is allowed free movement to spontaneously adjust their posture on the glass platform. The scene (Figure 3) shows the baby and the physiotherapist continuously, with other people appearing intermittently, in addition to the platform and the ceiling. The baby can be positioned anywhere on the platform, as can the physiotherapist and the other people in the scene. The lighting conditions were those of the room itself, without the introduction of any active or structured light sources. An attempt was made to avoid shadows, reflections, and glare from artificial lighting in the scene, although these occasionally occurred.

[0145] In several sessions with different physiotherapists, videos of 26 babies were captured at a sampling rate of 25 frames per second (25 fps), which is sufficient to capture key moments. The average duration of each video capture was approximately 2 minutes, with the baby spending roughly half the time in a prone position and the other half in a supine position. Each video contains approximately 3,000 frames. Once the video was captured, the individual frames, which are static images within a sequence of moving images, were automatically extracted. The extracted images (Figure 3), with a spatial resolution of 2704 x 2028 pixels and encoded in the RGB color space, are stored in JPEG format. From the 26 recorded videos, a set of approximately 76,000 images is available, a sufficient quantity to construct a dataset for semantic image segmentation.

[0146] For each frame, it is necessary to obtain the interpreted image corresponding to the segmentation, which constitutes the reference or ground truth about the actual content of an image. To this end, the following characteristic features or classes to be identified have been defined:

[0147] 1. Area outside the platform with no significant events.

[0148] 2. Metal surface of the platform.

[0149] 3. Internal area of ​​the platform corresponding to the glass.

[0150] 4. The evaluators and the baby's parents, if they appear in the image.

[0151] 5. The baby's body.

[0152] 6. The support that the baby provides on the glass surface.

[0153] The physiotherapists then proceeded to color and annotate these characteristic features in each of the frames. Their experience in creating the Ground Truth was fundamental, especially considering the enormous effort involved in creating a dataset of this size. In this way, the real image was transformed to obtain the corresponding interpreted image, generated by an expert annotator (Figure 4).

[0154] Each pixel in the Ground Truth image has been manually labeled by assigning it to one of six defined classes. Thus, each pixel in the Ground Truth image contains the RGB value corresponding to the color associated with the class in which the expert has classified it.

[0155] Specifically, 1 out of every 25 images has been colored. Therefore, the final dataset will consist of 3,000 image pairs, composed of the original image and its corresponding segmented image, which will be used in the training and validation phases of the CNN developed for this domain.

[0156] Example 2.2. Network selection and training

[0157] ResNet-50 (a 50-layer residual network) was selected as the deep convolutional neural network because it is one of the most powerful and widely used architectures in the field of data learning. The ResNet-50 implementation used had been pre-trained on the ImageNet dataset and served as the base model for transfer learning. The defined network consists of 206 layers, 227 connections, 1 input element, and 1 output element.

[0158] Six classes and images of size 2704 x 2028 pixels with 3 RGB channels were considered.

[0159] From the total set of images available in the own dataset, a 60 / 20 / 20 division was made, so that 60% of the images make up the training set, 20% the validation set and the remaining 20% ​​the test set.

[0160] During training, the evolution of accuracy and the loss function was monitored. The training process converged after 1,200 iterations, reaching an accuracy close to 1 and a loss function value close to 0, indicating very acceptable performance of the CNN model.

[0161] The training process was carried out on an NVIDIA GeForce RTX 2090 GPU with 11GB of RAM. The elapsed time was 4,258 minutes, which is equivalent to approximately 70 hours (2.8 days) of computation to complete 1,200 iterations. Once this process was finished, the trained network was exported to an executable file, which allows for multiclass semantic segmentation (six classes).

[0162] Example 3. RESULTS

[0163] Following the training process, a deep neural network composed of 206 layers was obtained, designed to automatically detect the support zones of babies in new input images. Each new image is processed in 6 seconds by the trained model. A qualitative and quantitative analysis of the results is presented below to evaluate the performance of the network trained on the custom dataset.

[0164] Qualitative analysis of network performance

[0165] For the qualitative analysis, two infants were randomly selected from the total set of 26. First, the frames annotated and segmented by the experts, which constituted the Ground Truth and were used to train the CNN, were qualitatively compared with the frames segmented by the CNN, which were the result of semantic segmentation. Figures 5 and 6 show a randomly selected sequence of frames, composed of three frames in a supine position and three in a prone position, corresponding to two different infants:

[0166] • The frame number appears in the first column

[0167] • In the second one, the actual original frame extracted from the video

[0168] • In the third image, the frame segmented by the experts, which is part of the Ground Truth with which CNN has trained.

[0169] • and in the fourth, the frame segmented by the CNN, where the original frame has been kept as a background to facilitate comparison with the real frame.

[0170] As can be seen (Figure 5 and Figure 6), CNN detects the following significant facts:

[0171] • In a prone position, identify the support areas in the trunk, abdomen and pelvis, as well as in the upper and lower limbs.

[0172] • In a supine position, identify the support areas on the head, back, pelvis and buttocks.

[0173] In the analyzed frames, the baby's support areas are clearly visible and consistently detected, regardless of the baby's position. Furthermore, the detection of these support areas was unaffected by factors such as diaper use, the baby's weight, or size. Additionally, the presence of multiple people in the room, different evaluators, and variations in lighting did not hinder the identification of the support areas and therefore did not influence the model's performance, indicating its generalizability and robustness.

[0174] Secondly, physiotherapy experts conducted a qualitative assessment of the results obtained by the CNN using frames that had never been shown to it during training. Figures 7 and 8 show the sequence of images corresponding to the two selected infants, used for the visual comparison of the results generated by the network.

[0175] • The frame number appears in the first column

[0176] • In the second column, the original frame extracted from the video for reference.

[0177] • in the third one, the frame segmented by CNN

[0178] • The opinion of the physiotherapists will be noted in the fourth column.

[0179] These two figures do not include frames from the Ground Truth because these frames were never colored by the experts. To facilitate the interpretation of the physiotherapists' comments, the following legend was established:

[0180] • In prone position: support on the face (AC), support of the trunk (AT), support of the abdomen (AA), support of the pelvis (AP), support of the upper limbs (AMMSS) and support of the lower limbs (AMMII).

[0181] • In supine position: head support (AC), back support (AE), buttock support (AG), upper limb support (AMMSS) and lower limb support (AMMII).

[0182] Comparing the frames obtained by the CNN with the expert criteria (Figure 7 and Figure 8) allows us to conclude that the CNN accurately detects most of the baby's body positions. Occasionally, some minor discrepancies have been observed, such as:

[0183] • CNN detects certain support levels with considerable accuracy, although it has not fully identified the area of ​​support according to expert criteria.

[0184] • In some cases, the CNN does not detect the supports corresponding to the lower and upper limbs.

[0185] • In some situations, the CNN has partially colored areas of the evaluator with colors corresponding to the glass or the baby. However, the percentage of pixels with incorrect colors is negligible.

[0186] • On other occasions, CNN detected head supports that had not initially been identified by the expert, but only after a more thorough analysis of the image.

[0187] Quantitative analysis of network performance

[0188] For the quantitative evaluation of the segmentation network's performance, the normalized confusion matrix (Figure 9) was calculated on the test set. This matrix represents, for each actual class (Ground Truth), the percentage of instances classified into each predicted class, with percentages normalized by row. This allows for an intuitive visualization of the network's overall behavior. In an ideal scenario, the predicted classes would exactly match the actual classes, so the cells on the main diagonal should approach 100%, while the remaining values ​​would tend towards 0%.

[0189] Thus, from Figure 9 it can be observed that:

[0190] • 91.8% of the pixels were predicted to belong to the Supported class and, indeed, they correspond to that class

[0191] • 7.0% of the pixels were classified as Supported class, although they actually belong to the Unsupported class, which labels the child's surfaces that are not supports.

[0192] • 90.7% of the pixels were predicted to belong to the Unsupported class and, indeed, they correspond to that class

[0193] • 8.0% of the pixels were classified as being of the Unsupported class, although they actually belong to the Supported class.

[0194] • 4% of the pixels were labeled as being of the Unsupported class when they are actually of the Person class, which labels physiotherapists and other people.

[0195] • The error percentage for the remaining classes is negligible, with values ​​close to 0%.

[0196] Analysis of the normalized confusion matrix shows that the network exhibits a high level of accuracy in class detection, with values ​​on the main diagonal exceeding 90%. Residual confusions occur primarily between supported and unsupported areas. This behavior can be attributed to the visual similarity of areas near the support surface, where differences in lighting, shadows, or pressure are minimal and difficult to distinguish, even for a human observer. Overall, the results indicate that the network has learned effectively and generally to differentiate the body's contact regions with the surface, successfully fulfilling the main objective.

[0197] The trained model has achieved an average accuracy of approximately 98%. This value implies that most pixels have been correctly classified into each class. However, because the classes are unbalanced—the Glass and Ceiling zones cover a larger area compared to the Supported and Unsupported zones—the accuracy metric ceases to be a reliable indicator, as it may overestimate the model's actual performance. To evaluate the reliability and accuracy of the trained CNN in segmenting pixels into the defined classes, a set of class-level metrics is used. The use of multiple metrics is common practice in machine learning tasks, as a model can perform well in one metric but be suboptimal in another. Furthermore, class-level metrics allow for a more specific analysis of each class's contribution and impact on the network's overall performance.

[0198] Table 1

[0199] Metrics obtained at the class level

[0200] From Table 1, it can be observed that:

[0201] • Regarding accuracy, the model achieves 82% in the Supported class, indicating that it correctly identifies most of the pixels where the baby is actually supported, although it may mark some areas as supported when they are not (false positives) according to Ground Truth. These differences may be due to shadows or areas of partial contact.

[0202] • The trained CNN exhibits a recall (sensitivity) greater than 91% in the Supported class, demonstrating a high capacity to correctly detect most of the pixels belonging to that class, with a reduced number of pixels not detected as supported (few false negatives).

[0203] • The Fl-score metric, which combines accuracy and sensitivity, reaches a value above 86% in the Supported class. This value reflects a good balance between the model's ability to identify all actual support zones (sensitivity) and its accuracy in correctly classifying these zones. This result confirms that the model consistently and reliably segments the contact areas. • In the loU metric, the Supported and Unsupported classes, which are the smallest, register the lowest values, although still above 76%. This result suggests that there may be annotation errors in Ground Truth in the Supported and Unsupported categories, where the boundaries between them may be blurred due to their proximity to the contact plane, making it difficult to precisely delineate both zones.

[0204] • Regarding the MeanBF Score metric, all classes show a high score, indicating good overall boundary alignment (Boundary Fl) between the Ground Truth image and the CNN-segmented image. However, the Supported and Unsupported classes show a lower value, although still above 80%, consistent with fuzzy boundary delimitation, as discussed previously.

[0205] Overall, the results obtained across the various evaluation metrics confirm that the trained CNN exhibits high and stable performance in segmenting the areas of the baby's body where it is supported and those where it is not. Furthermore, the model demonstrates high reliability and generalizability across all classes, making it suitable for this application domain.

[0206] Network check on a single image

[0207] To check the results of semantic segmentation on images not used in training, an individual image was randomly selected from the test set, specifically image 35. Figure 10 shows the corresponding results for that image.

[0208] In the first row, the original image, the CNN-segmented image, and both are superimposed, allowing for a direct visual comparison between the original and the CNN-segmented image. Qualitatively, a high degree of agreement is observed between the two images, indicating that the model is robust in identifying the six classes.

[0209] The second row displays the image segmented by experts (Ground Truth), the image segmented by CNN, and a composite image where both segmented images have been superimposed in different color bands. This composite image serves to show the differences between the Ground Truth image and the CNN result. The green and magenta regions highlight the areas where the CNN segmentation differs from Ground Truth. Visually, it can be observed that the image generated by CNN matches completely for the larger surface categories, such as Roof and Glass. In the smaller surface categories, such as Frame, Person, Supported, and Not Supported, small localized differences are visible, represented by green and magenta regions.

[0210] The amount of overlap between the two images is evaluated using the Jaccard index, also known as the Jaccard similarity coefficient or the Intersection-to-Union (IoU) ratio. This coefficient measures the similarity between two sets of pixels and is defined as the ratio between the intersection and the union of the corresponding regions in both images.

[0211] In this case, it is calculated, for each of the classes, as the intersection between the image segmented with the CNN and the Ground Truth image, divided by the union of both images.

[0212] The Jaccard coefficient calculated for image 35 of the test set is shown in Table 2. As can be seen, this metric confirms the results of the visual analysis: the highest Jaccard coefficient values ​​correspond to the Ceiling and Glass classes, while a lower value is obtained for Unsupported. Consequently, a greater overlap between the CNN-segmented image and the Ground Truth image is confirmed for the classes with the largest surface area, and the slight discrepancies in the Supported and Unsupported classes can be attributed to the difficulty of precisely defining the limits of the baby's support.

[0213] Table 2 Jaccard index of image 35 Importance of supports according to the age of the baby

[0214] Through the evaluation of spontaneous motor skills, it has been observed that the support and straightening points of the scapula and pelvis, that is, of the shoulder and pelvic girdle, were less competent in babies who were around 3 months old than in babies of

[0215] 6 months of age, in which support and straightening were more effective. Taking into account the support and straightening of the shoulder and pelvic girdle, the results of the frames processed by the CNN model, in supine position, confirm that babies under 3 months of age presented less support area (green zone) at the caudal level, while in babies close to 6 months of age, the green support area was wider at the caudal level (Figure 11).

Claims

CLAIMS 1. Device (1) configured for the acquisition of images and / or videos of the support points and areas of a subject (B) on a surface, comprising: a structure (2) formed by a rectangular frame provided with support elements in contact with the ground, a transparent support (3) disposed on the rectangular frame on which the subject would be placed, characterized in that the device (1) comprises a support (5) for a camera (6) located at the intersection of the support elements of the rectangular frame with the ground and below the transparent support (3).

2. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to claim 1, characterized in that it comprises stops (4) through which the transparent support (3) rests on the rectangular frame.

3. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that each support point of the rectangular frame with the ground forms a right angle, so that it has a section that extends horizontally along the ground and another section that rises vertically to meet the rectangular frame.

4. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that it comprises two support points of the rectangular frame with the ground, each of them in the shape of a U, or four support points of the rectangular frame with the ground, each of them in the shape of an L.

5. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that the rectangular frame is square.

6. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that the subject is a person, preferably a baby or a child.

7. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that the subject is placed on the transparent support (3) in a dorsal and ventral decubitus position.

8. Device (1) configured for the acquisition of images and / or videos on the support points and areas of a subject (B) on a surface, according to any of the preceding claims, characterized in that the transparent support (3) has corner protectors (7).

9. Device, according to any of claims 1 to 8, characterized in that the transparent support (3) is made of vitreous material or equivalent optically transparent material.

10. Device, according to any of claims 1 to 9, characterized in that it is configured for the continuous capture of high-resolution video or successive images.

11. Device, according to any of claims 1 to 10, characterized in that the acquired images have sufficient resolution to identify the subject's support areas and are stored in a standard digital format.

12. Device, according to any of claims 1 to 11, characterized in that the image acquisition is carried out under ambient lighting conditions without requiring active or controlled light sources.

13. Device, according to any of claims 1 to 12, characterized in that the subject is an infant or young child placed on the transparent support (3).

14. Device, according to any of claims 1 to 13, characterized in that the support detection is robust against variations in lighting conditions, presence of accessories or anthropometric differences of the subject.

15. System for automatically identifying the support zones of a subject (B) placed on a transparent support (3), comprising the device according to any of the preceding claims and a processor configured to perform image segmentation into multiple classes or regions of interest.

16. System, according to claim 15, characterized in that the processor uses a convolutional neural network model trained by transfer learning.

17. System, according to any of claims 15 or 16, characterized in that the trained model achieves a level of precision sufficient to distinguish the support zones of the subject (B) with high reliability.

18. System, according to any of claims 15 or 17, characterized in that the model is configured to distinguish at least between contact regions and non-contact regions of the subject (B) with the transparent support.

19. A method for generating a dataset for training a machine learning model to segment images of subjects on a transparent support, comprising: (i) capturing video sequences or images of the subject (B); (ii) extracting individual frames or samples; (iii) labeling the images by identifying the regions corresponding to different classes; and (iv) using said images to train, validate, and test the model.

20. Method, according to claim 19, characterized in that the annotation of the images is performed manually or semi-automatically by expert personnel.

21. A method, according to any of claims 19 or 20, characterized in that the model training is carried out until convergence or stability of the learning parameters is achieved.

22. Use of the system of claims 15 to 18 to identify the support zones of a subject (B) in different positions on the transparent support (3).

23. Use of the system of claims 15 to 18 to compare the automatically identified support zones with those evaluated by a qualified professional.

24. Computer program containing instructions executable by a processor to perform image or video segmentation according to the procedure of claims 19 to 21.