Apparatus and method for processing image data

The apparatus and method analyze image data using machine-learning models to detect user and environmental conditions, providing timely notifications for improved safety in AR/VR devices.

WO2026159042A1PCT designated stage Publication Date: 2026-07-30SONY SEMICON SOLUTIONS CORP +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONY SEMICON SOLUTIONS CORP
Filing Date
2026-01-20
Publication Date
2026-07-30

AI Technical Summary

Technical Problem

Existing AR/VR technologies lack the capability to effectively analyze user and environmental conditions in real-time to provide timely notifications for potential risks or anomalies, particularly in relation to user behavior and ground conditions, which can lead to accidents.

Method used

An apparatus and method that utilize processing circuitry to analyze image data from a top-down perspective, employing machine-learning models to detect user conditions and environmental hazards, and trigger notifications through integrated sensors for enhanced safety.

Benefits of technology

Enables real-time detection and notification of user risks and environmental hazards, improving user safety by predicting potential accidents and health issues, thus enhancing the usability of AR/VR devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2026051230_30072026_PF_FP_ABST
    Figure EP2026051230_30072026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an apparatus for processing image data. The apparatus includes processing circuitry configured to receive the image data indicating a region of the ground in front of a user and a portion of the user's lower body from a top-down perspective relative to the user. The processing circuitry is further configured to determine, based on the image data, at least one of a condition related to the user and a condition related to the user's environment. The processing circuitry is further configured to cause notification of the user if the condition is determined to present a risk or an anomaly.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Apparatus and method for processing image data

[0002] Field

[0003] The present disclosure relates to an apparatus and a method for processing image data.

[0004] Background

[0005] An Augmented Reality (AR) headset may relate to a wearable device that overlays digital content onto the real world, enhancing the user’s perception with interactive holograms, visual cues, or contextual information. It may employ transparent or semi-transparent displays, along with cameras, depth sensors, and motion tracking to integrate virtual elements with the physical environment in real time.

[0006] A Virtual Reality (VR) headset may relate to a fully immersive device that replaces the user’s real-world view with a computer-generated 3D environment, creating the illusion of presence within a simulated space. Unlike AR headsets, VR systems may use opaque displays and head-tracking sensors to provide an interactive, enclosed experience, often requiring additional controllers for navigation and interaction.

[0007] There may be a demand for improved AR / VR technologies.

[0008] Summary

[0009] This demand may be satisfied by the subject-matter of the independent claims. Further aspects are set forth in the dependent claims, the drawings, and the following description.

[0010] According to a first aspect, the disclosure provides an apparatus for processing image data. The apparatus includes processing circuitry configured to receive the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user. The processing circuitry is further configured to determine, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment. The processing circuitry is further configured to cause notification of the user if the condition is determined to present a risk or an anomaly.According to a second aspect, the disclosure provides a method for processing image data. The method includes receiving the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user. The method further includes determining, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment. The method further includes causing notification of the user if the condition is determined to present a risk or an anomaly.

[0011] According to a third aspect, the disclosure provides a non-transitory machine-readable medium comprising instructions, which, when the instructions are carried out on an apparatus, cause the apparatus to carry out the method according to the second aspect.

[0012] Brief description of the Figures

[0013] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which

[0014] Fig. 1 depicts a block diagram of an apparatus according to the present disclosure;

[0015] Fig. 2 depicts a flowchart of a method according to the present disclosure; and

[0016] Fig. 3 depicts a use-case of the present disclosure.

[0017] Detailed Description

[0018] Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.

[0019] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.

[0020] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.

[0021] Fig. 1 depicts an apparatus 100 for processing image data. An image may refer to a structured representation of spatial, visual, or multidimensional information, e.g., encoded as a function of intensity, color, depth, position, or other measurable properties across a given domain. It may be acquired and optionally processed through various modalities, including optical sensors (e.g., CCD (charge coupled device), CMOS (complementary metal oxide semiconductor), LiDAR, time-of-flight, radar, infrared imaging), computational techniques (e.g., ray tracing, neural networks). Depending on the application, images may exist as 2D pixel arrays, 3D point clouds, volumetric reconstructions, or hyperspectral datasets, and may be stored in raster formats (e.g., PNG (portable network graphics), JPEG (joint photographic experts group)), vector formats (e.g., SVG (scalable vector graphics)), specialized formats (e.g., PLY (polygon file format)), or the like. Advanced imaging techniques, such as time-of-flight (ToF), structured light, and synthetic aperture imaging, may extend the concept of two-dimensional imaging beyond static intensity maps to capture depth, motion, and material properties. More generally, an image may be any representation of information that encodes spatial, spectral, or temporal attributes, applicable in fields such as computer vision, medical imaging, remote sensing, and artificial intelligence.

[0022] Accordingly, image data may refer to any type of data to represent an image. The image data may be based on raw or processed numerical values corresponding to pixel intensity, color channels, depth measurements, or other relevant features, stored in formats such as raster arrays, point clouds, or volumetric datasets. Depending on the imaging technique, image datamay include metadata such as resolution, sensor parameters, or timestamp information, enabling analysis, processing, and interpretation in applications like computer vision, medical imaging, and remote sensing.

[0023] The apparatus 100 includes processing circuitry 110. For example, the processing circuitry 110 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC) a neu-romorphic processor or afield programmable gate array (FPGA). The processing circuitry 110 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 100 may include memory configured to store instructions, which when executed by the processing circuitry 110, cause the processing circuitry 110 to perform the steps and methods described herein.

[0024] The processing circuitry 110 configured to receive the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user.

[0025] The image data may be acquired by a corresponding image sensor which may be included in the apparatus 100, such that the circuitry may receive the image data directly or indirectly from the sensor. For example, the image sensor may be (part of) a camera of a head-mounted device (as the apparatus 100) which faces downwards (relative to a front-camera when a user wears the head-mounted device) to capture the portion of the user’s lower body from the top-down perspective. A head-mounted device may include smart glasses, a smart headset, virtual reality glasses / headset, augmented reality glasses / headset, or the like. However, the present disclosure is not limited to the case that the apparatus 100 includes the image sensor. For example, the apparatus 100 may be (included in) a server or a cloud and may thus receive the image data from a different device that may include the image sensor.

[0026] Accordingly, when the device is worn in a typical position (i.e. , eyepieces over the eyes), a camera provided for the apparatus may be facing downwards and thus, the image data may indicate a region of the ground in front of the user as well as at least a portion of the user’s lower body from a top-down perspective. For example, the principles of the present disclosure may be applied while the user is walking and the image data may be used to analyze the user’s walking behavior and the ground in front of the user, for example to determine whetherthe user should be cautious about where they are stepping (but the present disclosure is not limited in that regard). The portion of the user’s lower body may, for example, include at least one of a foot of the user (or both feet), a lower leg (or both lower legs), an upper leg (or both upper legs), a hip region, a knee region, or the like.

[0027] The circuitry 110 is further configured to determine, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment.

[0028] For example, the image data may be fed into an algorithm configured to determine the condition, such as a trained machine-learning model / algorithm. A (trained) machine-learning (or machine-learned) model may refer to a data structure and / or set of rules representing a statistical model that the processing circuitry 110 may use to generate at least one machinelearning output. It should be noted that the present disclosure may be carried out with one machine-learning model that is configured to determine both the condition related to the user and the condition related to the user’s environment. Alternatively, each condition may be determined with a separate machine-learning model, in some examples. The data structure and / or set of rules may represent learned knowledge (e.g. based on training performed by a machine-learning algorithm as described below). In machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.

[0029] The machine-learning model may be trained based on a training algorithm. The term "training algorithm" may denote a set of instructions that are used to create, train or use a machinelearning model. For the machine-learning model to determine an output, the machine-learning model may be trained using training data, as commonly known. By training the machinelearning model with a large set of training data and associated training information, the machine-learning model may learn to determine the output. By training the machine-learning model using training information (e.g., training image data), the machine-learning model may learn a transformation between training input data and appropriate output data.

[0030] The machine-learning model may be trained using training input data (e.g. training image data). For example, the machine-learning model may be trained using a training method called "supervised learning". In supervised learning, the machine-learning model may be trained using a plurality of training samples, wherein each sample may include a plurality of input data values, and a plurality of desired output values, i.e. , each training sample may be associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model may learn which output value to provide based on aninput sample that is similar to the samples provided during the training. For example, a training sample may include input data that is based on training image data and desired condition relating to the user and / or the user’s environment.

[0031] Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples may lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained machine-learning model may be restricted to a limited set of values (categorical variables), i.e. , the input is classified to one of the limited set of values. Similarity learning algorithms may be similar to classification algorithms but may be based on learning from examples using a similarity function that measures how similar two related objects are.

[0032] Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data may be supplied and an unsupervised learning algorithm may be used to find structure in the input data (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data including a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.

[0033] Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") may be trained to take actions in an environment. Based on the taken actions, a reward may be calculated. Reinforcement learning may be based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).

[0034] Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and / or the machine-learning algorithm may include a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step beforeperforming classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example.

[0035] For example, the machine-learning model may be an Artificial Neural Network (ANN). An ANN may refer to a system that is inspired by biological neural networks, such as in retina or in brains. ANNs may include a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There may be three types of nodes: Input nodes that receive input values (e.g., training image data), hidden nodes that are (only) connected to other nodes, and output nodes that provide output values (e.g., condition relating to the user and / or user’s environment). Each node may represent an artificial neuron. Each edge may transmit information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and / or of edges may be adjusted in the learning process. In other words, the training of an ANN may include adjusting the weights of the nodes and / or edges of the ANN, i.e., to achieve a desired output for a given input.

[0036] Alternatively, the machine-learning model may include a different structure and, e.g., be a support vector machine, a random forest model or a gradient boosting model. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.

[0037] In examples, the machine-learning model may be a combination of any of above examples.

[0038] However, the present disclosure is not limited to using a machine-learning model since static algorithms, statistic algorithms, rule-based algorithms, or the like may be employed, in some examples.

[0039] The processing circuitry 110 is further configured to cause notification of the user if the condition is determined to present a risk or an anomaly. The notification may be any type of message that is output to the user to notify them of a certain circumstance (i.e., the risk or anomaly), such as a sound, a visual notification (e.g., on a display), a haptic notification (e.g, vibration), or the like.

[0040] For example, the condition related to the user may include a user behavior, such as a walking behavior or a movement pattern. A walking behavior may refer to a manner in which the user moves on foot, influences by biomechanics, environmental factors, cognitive contexts, socialcontexts, and the like. A movement pattern may refer to a sequence and rhythm of movements during locomotion e.g., involving alternating steps. Movement pattern and walking behavior may be used interchangeable in the context of the present disclosure. For example, a possible / predicted walking path, future foot position (or step position), walking pace, stopping time, a walking direction, a possible destination, or the like may be determined. The determining of the user behavior may be based on at least one of a walking rhythm, a stride length, and a foot position, or the like. For example, such parameters may be obtained based on the image data by analyzing at least one foot of the user. An exemplary approach to determine the user behavior may include to assume that the walking rhythm and the stride length may be (roughly) constant for a predetermined amount of time and to extrapolate the user’s behavior within that predetermined amount of time. However, more complex algorithms may applied which may take into account other parameters, such as a road map and obstacles in the way of the user such that a possible detour of the user may be considered / determined. In some examples, the processing circuitry 110 is configured to predict a future user behavior based on a machine-learning model using the image data, as discussed above.

[0041] For example, if a step prediction (or determination of future foot position) is carried out, a dedicated image sensor with at least one embedded machine-learning model (“intelligent vision sensor”) may be used (e.g., Sony IMX500, but the present disclosure is not limited in that regard). In such a sensor, Al processing capabilities may be provided directly within the hardware, thereby allowing real-time analysis of the user's gait and movement patterns on the apparatus. Such embedded machine learning models may monitor the user's walking rhythm, stride length, and foot positioning by analyzing visual data from the top-down perspective. This may aid in determining typical movement patterns without the need for external data processing. For example, trajectory estimation may be utilized. In such an approach, the user's next steps may be predicted in real-time. By incorporating video prediction techniques, the machine learning model may analyze sequences of visual data to forecast future movements and foot placements (as the future user behavior, but the present disclosure is not limited in that regard since also a future path may be predicted, in some examples). Video prediction may allow the system to generate future frames based on past and current visual inputs, thereby anticipating and projecting the user's motion. This may enable the system to determine an area where the user's foot is likely to land. By performing computations directly on the image sensor, the system may minimize data transmission requirements, leading to faster response times and enhanced user privacy. This predictive capability may allow for timely detection of hazards before the user steps on them, providing an opportunity to warn the user and prevent potential accidents. The system may process visual data instantly on the sensor's Al module, classifying hazards without delay and eliminating the need for datato be sent to external processors. Sensitive visual information may not be needed to be stored or transmitted elsewhere, ensuring user and bystander privacy. For example, when a user goes to a toilet, the circuitry 110 may process all visual data locally, ensuring that no sensitive information is captured or transmitted, thus maintaining the privacy of the user in sensitive situations. Moreover, by predicting the user’s steps, increased safety for users may be achieved, e.g., from slippery steps or from stepping on nails (e.g., by notifying the user of the risk). Movement safety of visually impaired people may further be enhanced.

[0042] In some examples, the condition related to the user may include a lower body anomaly. A lower body anomaly may refer to any deviation from typical anatomical structure, function or behavior in the lower extremities of the user, including abnormalities in bones, muscles, joints, or nerves that may affect movement or stability. The lower body anomaly may also be indicative of an underlying condition that may cause the deviation. For example, the user’s lower body may be monitored to detect signs of incorrect or unhealthy gait that may lead to or be based on poor / bad / improper posture or indicate underlying health issues. As indicated above, the image data may indicate real-time visual data of at least one of the user's legs and the user’s feet as they move. Embedded machine learning algorithms may analyze this data to assess the user's gait patterns continuously. By observing parameters such as stride length, walking speed, foot placement, and joint movement, the processing circuitry 110 may be configured to identify deviations from normal / healthy walking patterns. Subtle anomalies like limping, uneven weight distribution, or irregular step timing may be detected by comparing current movements to the user's established baseline gait or to a known healthy baseline gait. These anomalies may signal issues such as muscle weakness, joint problems, balance disorders, early signs of neurological conditions, or the like. If such an anomaly is detected, immediate feedback may be provided to the user, e.g., through a display or other alert mechanisms (such as sound). In other words, the notification may be caused based on the lower body anomaly, and the notification may indicate a correction of at least one of the user’s walking pattern and the user’s posture based on the lower body anomaly. Such an approach may allow the user to make conscious adjustments to their posture or seek professional medical advice before minor issues develop into more serious conditions. Overtime, the system may track changes in gait patterns, offering valuable insights into the user's musculoskeletal health and helping to prevent long-term complications associated with improper gait. In more general terms, health incidents may be detected quickly.

[0043] In some examples, the condition related to the user’s environments may include a condition of the region of the ground in front of the user. The condition of the region of the ground in front of the user may refer to any physical characteristic of the surface, such as its texture,slope, stability, and presence of obstacles, which may influence gait, balance, and traction. The processing circuitry 110 may be configured to scan the predicted step area for potential hazards (e.g., based on a machine-learning model as discussed above). Thereby, risks may be identified, such as:

[0044] • Slippery Surfaces: Substances may be detected that reduce traction, such as water, ice, or oil spills by recognizing specific visual patterns and textures.

[0045] • Unstable Ground: Uneven surfaces, loose gravel, potholes, or shifting terrain may be identified by analyzing depth cues and surface irregularities.

[0046] • Sharp Objects: Harmful debris may be identified, such as nails, glass shards, or rocks through, e.g., based on object detection algorithms.

[0047] • Obstacles: Curbs, steps, or unexpected changes in elevation may be detected that may pose tripping hazards.

[0048] A corresponding alert mechanism may be provided for notifying the user to give immediate feedback to the user, e.g., through visual cues on the AR display, auditory alerts, or haptic feedback (e.g., vibrations), allowing for quick reaction to potential dangers.

[0049] In other words, in some examples, the processing circuitry 110 may be further configured to assess a risk for the user based on the condition of the region of the ground in front of the user and based on the predicted walking behavior. In such examples, the notification may be an alert to the user in accordance with the risk. For example, depending on a severity the risk, different alerts may be issued, e.g., which may differ in alert intensity, alert medium (e.g., auditory, visual), or the like.

[0050] Hence, in such an example, a combination of step prediction and danger recognition may be employed to enhance the user's safety by reducing the risk of accidents through timely detection and alerts about potential ground hazards.

[0051] Fig. 2 depicts a flowchart of a method 200 for processing image data according to the present disclosure. The method 200 includes receiving, 210, image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user. The method 200 further includes determining, 220, based on the image data, at least one of a condition related to the user and to the user’s environment. The method 200 further includes causing, 230, a notification of the user if the condition is determined to present a risk or an anomaly.Fig. 3 depicts an exemplary use case of a head-mounted assistive vision device 300 (as an apparatus according to the present disclosure). Fig. 3 depicts an elderly woman (user) 310 walking along an urban sidewalk while wearing the head-mounted assistive vision device 300. In this example, the device is designed for visually impaired users and includes a structured headset equipped with multiple sensors for real-time environmental scanning. The device 300 may employ optical and depth-sensing technologies to detect objects and surface irregularities. The woman 310 is approaching an area where broken glass is scattered on the pavement, presenting a potential hazard that may not be easily detectable through conventional mobility aids. The device 300 may be configured to emit structured light (or an infrared scanning beam, in some examples), thus using an active sensing mechanism that may acquire spatial data in real time. The device 300 may be configured to process the acquired data using embedded computational units that carry out at least one of the methods discussed herein. Based on the analysis, the device 300 may be able to distinguish between walkable surfaces and obstacles such as debris, curbs, or unexpected obstructions. Upon detection of a potential hazard, the system may trigger an alert through auditory cues, haptic feedback, or visual overlays (if equipped with augmented reality capabilities). The device 300 may be configured to handle variable lighting conditions, motion dynamics, and diverse obstacle types. The sensors may be capable of recognizing objects under different illumination levels, detecting reflective materials like glass, and differentiating between transient and static hazards. Additionally, real-time depth estimation may enable the system to prioritize obstacles based on proximity and risk level. This assistive vision system exemplifies a wearable mobility enhancement tool designed to provide visually impaired users with increased spatial awareness. By integrating multi-sensor fusion, real-time processing, and adaptive alert mechanisms, the device 300 may offer a proactive solution for safe and independent navigation in complex environments.

[0052] According to the present disclosure, an egocentric camera system mounted to an ARA / R headset may be provided. The camera system may be positioned top-down facing. This direction may ensure privacy preservation by not filming or analyzing other people, car number plates, or the like, and for functional reasons. The camera system may record the area on the ground in front of the user and at least a part of the user’s lower body. The recorded footage (image data) may be analyzed by the system’s onboard processor or in a cloud. The footage may be discarded afterwards or stored. The system may predict a position of a next step and analyze an area on the ground for potential dangers such as slippery ground, unstable ground, nails, glass shards, or the like. Further, the system may analyze the user’s lower body for anomalies that could be in connection with health hazards. For example, in addition to the above-mentioned health issues, the user may be having a seizure such that their body mightbe shaking. The system may detect such anomalous shaking and call for help. The user may also be in contact with electric wire which may be detected as well or might have lost conscience under water or in space.

[0053] More details and aspects of the methods described herein are explained in connection with the proposed technique or with one or more examples described above (e.g., Figs. 1 to 3). The methods may include one or more additional optional features corresponding to one or more aspects of the proposed technique or to one or more examples described above.

[0054] The following examples pertain to further embodiments of the present disclosure:

[0055] (1) An apparatus for processing image data. The apparatus includes processing circuitry configured to receive the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user. The processing circuitry is further configured to determine, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment. The processing circuitry is further configured to cause notification of the user if the condition is determined to present a risk or an anomaly.

[0056] (2) The apparatus of (1), wherein the condition related to the user comprises a user behavior.

[0057] (3) The apparatus of (2), wherein the determining of the user behavior is based on at least one of a walking rhythm, a stride length, and a foot position.

[0058] (4) The apparatus of (2) or (3), wherein the determining of the user behavior comprises determining a movement pattern.

[0059] (5) The apparatus of any one of (2) to (4), wherein the processing circuitry is further configured to predict a future user behavior based on a machine-learning model using the image data.

[0060] (6) The apparatus of any one of (1) to (5), wherein the condition related to the user comprises a lower body anomaly.(7) The apparatus of (6), wherein the notification is caused based on the lower body anomaly. In such examples, the notification indicates a correction of at least one of the user’s walking pattern and the user’s posture based on the lower body anomaly.

[0061] (8) The apparatus of any one of (1) to (7), wherein the condition related to the user’s environments comprises a condition of the region of the ground in front of the user.

[0062] (9) The apparatus of (8), wherein the condition of the region of the ground in front of the user is determined based on a machine-learning model.

[0063] (10) The apparatus of (9), wherein the machine-learning model is further configured to predict a walking behavior of the user.

[0064] (11) The apparatus of (10), wherein the circuitry is further configured to assess a risk for the user based on the condition of the region of the ground in front of the user and based on the predicted walking behavior. In such examples, the notification is an alert to the user in accordance with the risk.

[0065] (12) A method for processing image data. The method includes receiving the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user. The method further includes determining, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment. The method further includes causing notification of the user if the condition is determined to present a risk or an anomaly.

[0066] (13) The method of (12), wherein the condition related to the user comprises a user behavior.

[0067] (14) The method of (13), wherein the determining of the user behavior is based on at least one of a walking rhythm, a stride length, and a foot position.

[0068] (15) The method of (13) or (14), wherein the determining of the user behavior comprises determining a movement pattern.

[0069] (16) The method of any one of (13) to (15), further including predicting a future user behavior based on a machine-learning model using the image data.(17) The method of any one of (12) to (16), wherein the condition related to the user comprises a lower body anomaly.

[0070] (18) The method of claim (17), wherein the notification is caused based on the lower body anomaly. In such examples, the notification indicates a correction of at least one of the user’s walking pattern and the user’s posture based on the lower body anomaly.

[0071] (19) The method of any one of (12) to (18), wherein the condition related to the user’s environments comprises a condition of the region of the ground in front of the user.

[0072] (20) The method of (19), wherein the condition of the region of the ground in front of the user is determined based on a machine-learning model.

[0073] (21) The method of (20), wherein the machine-learning model is further configured to predict a walking behavior of the user.

[0074] (22) The method of (21), further including assessing a risk for the user based on the condition of the region of the ground in front of the user and based on the predicted walking behavior. In such examples, the notification is an alert to the user in accordance with the risk.

[0075] (23) A non-transitory machine-readable medium comprising instructions, which, when the instructions are carried out on an apparatus, cause the apparatus to carry out the method of any one of (12) to (22).

[0076] (24) A computer program comprising instructions which, when the computer program is executed out on a computer, causes the computer to carry out the method of any one of (12) to (22).

[0077] The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.

[0078] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmedcomputers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and / or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.

[0079] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, -functions, -processes or -operations.

[0080] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.

[0081] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.

Claims

ClaimsWhat is claimed is:

1. An apparatus for processing image data, the apparatus comprising processing circuitry configured to:receive the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user;determine, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment; andcause notification of the user if the condition is determined to present a risk or an anomaly.

2. The apparatus of claim 1 , wherein the condition related to the user comprises a user behavior.

3. The apparatus of claim 2, wherein the determining of the user behavior is based on at least one of a walking rhythm, a stride length, and a foot position.

4. The apparatus of claim 2, wherein the determining of the user behavior comprises determining a movement pattern.

5. The apparatus of claim 2, wherein the processing circuitry is further configured to: predict a future user behavior based on a machine-learning model using the image data.

6. The apparatus of claim 1 , wherein the condition related to the user comprises a lower body anomaly.

7. The apparatus of claim 6, wherein the notification is caused based on the lower body anomaly, and wherein the notification indicates a correction of at least one of the user’s walking pattern and the user’s posture based on the lower body anomaly.

8. The apparatus of claim 1 , wherein the condition related to the user’s environments comprises a condition of the region of the ground in front of the user.

9. The apparatus of claim 8, wherein the condition of the region of the ground in front of the user is determined based on a machine-learning model.

10. The apparatus of claim 9, wherein the machine-learning model is further configured to predict a walking behavior of the user.

11. The apparatus of claim 10, wherein the circuitry is further configured to:assess a risk for the user based on the condition of the region of the ground in front of the user and based on the predicted walking behavior, wherein the notification is an alert to the user in accordance with the risk.

12. A method for processing image data, the method comprising:receiving the image data indicating a region of the ground in front of a user and a portion of the user’s lower body from a top-down perspective relative to the user;determining, based on the image data, at least one of a condition related to the user and a condition related to the user’s environment; andcausing notification of the user if the condition is determined to present a risk or an anomaly.

13. The method of claim 12, wherein the condition related to the user comprises a user behavior.

14. The method of claim 13, wherein the determining of the user behavior is based on at least one of a walking rhythm, a stride length, and a foot position.

15. The method of claim 13, wherein the determining of the user behavior comprises determining a movement pattern.

16. The method of claim 13, further comprising:predicting a future user behavior based on a machine-learning model using the image data.

17. The method of claim 12, wherein the condition related to the user comprises a lower body anomaly.

18. The method of claim 17, wherein the notification is caused based on the lower body anomaly, and wherein the notification indicates a correction of at least one of the user’s walking pattern and the user’s posture based on the lower body anomaly.

19. The method of claim 12, wherein the condition related to the user’s environments comprises a condition of the region of the ground in front of the user.

20. A non-transitory machine-readable medium comprising instructions, which, when the instructions are carried out on an apparatus, cause the apparatus to carry out the method of claim 12.