System for context-sensitive environmental perception with autonomous physical orientation control

The system autonomously controls sensor unit orientation based on interaction context, addressing the lack of context-aware orientation in existing systems by using a state machine and learning module for proactive and user-friendly state transitions.

DE202026001915U1Undetermined Publication Date: 2026-07-02BG SYSTEM GMBH

Patent Information

Authority / Receiving Office
DE · DE
Patent Type
Utility models
Current Assignee / Owner
BG SYSTEM GMBH
Filing Date
2026-04-28
Publication Date
2026-07-02

AI Technical Summary

Technical Problem

Existing systems for environmental perception lack the ability to autonomously control the physical orientation of sensor units based on the context of ongoing human-system interactions, failing to switch between semantically different perceptual states without manual user triggers.

Method used

A system comprising a sensor unit, actuator, and processing unit with a state machine that autonomously switches between perceptual states based on context information derived from ongoing human-system interactions, using a learning-based processing module and persistent context memory to control the actuator for spatial orientation adjustments.

Benefits of technology

Enables proactive and context-aware orientation control of sensor units, allowing for semantically distinct perceptual states without manual user activation, local data processing, and user-adapted interaction initiation, while ensuring data privacy and providing understandable state transitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

- Main claim (product claim) System for context-sensitive environmental perception, comprising a sensor unit for capturing environmental information, at least one actuator for physically changing the spatial orientation of the sensor unit, and a processing unit with a state machine, characterized in that the state machine switches autonomously between at least two semantically different perception states without manual triggering by a user, wherein the processing unit derives the currently active perception state from context information generated during an ongoing human-system interaction and controls the actuator in such a way that the sensor unit is aligned with a spatial area assigned to the active perception state.
Need to check novelty before this filing date? Find Prior Art

Description

1. Technical field of the invention The invention relates to a system for context-sensitive environmental perception, comprising a sensor unit, an actuator for physically changing the spatial orientation of the sensor unit, and a processing unit with a state machine. In particular, the invention relates to a system in which the state machine switches autonomously between semantically different perceptual states without manual triggering by a user, thereby controlling the physical orientation of the sensor unit. The system belongs to the technical field of human-machine interaction, in particular to the area of ​​interactive stationary systems with physical environment perception and autonomous behavior control. Exemplary application areas include interactive toys, personal assistants, smart home devices, and care assistance systems. The invention is not limited to a specific application. 2. State of the art Various systems are known in the prior art that use sensor units to acquire environmental information and sometimes combine these with actuators to change the spatial orientation. However, the known solutions each have shortcomings. 2.1 Reactive face tracking mechanisms Animatronic mechanisms (Nilheim Eye Mechanism, Will Cogley, 2019) and robotic systems (Watchman Robot, Graham Jessup) are known examples, in which a sensor unit is aligned with a detected face using actuators. These systems use facial recognition algorithms to determine the position of a face in the field of view and to track the sensor unit accordingly. Deficiencies: These systems operate exclusively reactively – they follow a visually recognized face without considering the context of an ongoing human-system interaction. There is only one operating mode (face tracking). Autonomous switching between semantically different perceptual states does not occur. The system cannot independently decide to redirect the sensor unit from a user to an object referenced in conversation. 2.2 Social robots with context-based interaction From US patent 11,148,296 B2 (Breazeal et al.) a persistent companion device is known that recognizes human interaction cues by analyzing various sensory inputs. The system has several movable segments that can assume poses corresponding to the context of the recognized interaction cue. It can perform emotion recognition, person tracking, and speech recognition, and adapt personality traits through repeated interactions. Furthermore, social robots are known from US 12,136,431 B2 (Embodied Inc.) that analyze multimodal inputs and track the engagement of multiple users to determine which users interact with the device. Furthermore, humanoid social robots of the Pepper and NAO types (SoftBank Robotics / Aldebaran Robotics) are known, which feature movable heads with built-in cameras and actuators, as well as systems for emotion recognition and proactive interaction control. Their patent portfolio (including WO 2013 / 178741 and WO 2015 / 158876) includes systems for behavior planning and dialogue control. Deficiencies: In these social robots, head movement serves for expression – social communication and reactive face tracking (output-side motion control). Input-side control – in which the physical orientation of a sensor unit autonomously switches between semantically different perceptual states, with each state assigned its own spatial area for environmental perception – is not claimed in any of the reviewed documents. 2.3 Reactive voice-controlled alignment systems From EP 3 894 972 A1 (Google LLC) a motorized device is known that autonomously adjusts its orientation. The system uses a microphone to detect spoken language, determines the user's position relative to the device, and controls motors to align a display with the user. Deficiencies: The system is triggered by the user's spoken utterance – it is a manual trigger in the form of a voice command. Without such an utterance, the system remains in its current state. It lacks a state machine that autonomously and proactively switches between different perceptual states based on the accumulated interaction context. 2.4 Reactive eye-tracking systems From EP 3 691 766 B1 (Sony Interactive Entertainment Inc.) a companion robot is known that detects a user's gaze direction and adjusts its position accordingly. The system tracks the current gaze direction and directs a projector onto a surface towards which the user is looking. Deficiencies: The system reacts to the current gaze direction. It does not accumulate contextual information about the course of an interaction and does not have semantically distinct perceptual states. The spatial area is determined directly by the gaze direction, not by a state machine. 2.5 Cloud-dependent interactive toy systems US patent 12,197,409 B2 (Me My Essence LLC) discloses an AI-powered interactive toy that captures stimuli from a user and transmits them to an external cloud server for processing. Shortcomings: The system requires a permanent network connection and transmits personal data to external servers. Physical control of the sensor unit's orientation via an actuator, based on the interaction context, is not provided. 2.6 Reactive voice assistants Commercially available voice assistants (Amazon Alexa, Google Nest) are stationary devices that wait for an activation word and then process voice commands. From EP 4 143 675 A1 (Google LLC) a conditional camera control by a voice assistant is known. The system receives a spoken utterance, determines whether a condition is met, and triggers a camera recording if the condition is met. Drawbacks: These systems are fundamentally reactive – every interaction is triggered manually by the user. There is no physical alignment of a sensor unit with different areas of the room based on the interaction context. 2.7 Cognitive robot architectures and proactive planning systems Cognitive robot architectures combining state machines and memory systems are known from the state of the art. The ArmarX framework (Vahrenkamp et al., 2015; Wächter et al., 2016; Peller-Konrad et al., 2023) implements statecharts (event-driven state machines), sensor-actuator units for camera components, and a hierarchical memory system with episodic interaction memory for humanoid robots. The published scenarios are triggered by user commands. Furthermore, EP 4 578 604 A1 (Honda Motor Co. Ltd., 2023) discloses a proactive robot system that uses theory-of-mind models and emotion-aware planning for autonomous behavior selection in human-robot interactions. Deficiencies: The ArmarX framework does not describe an application that autonomously aligns a sensor unit between semantically different perceptual states – all published scenarios require explicit user activation. The system according to EP 4 578 604 A1 operates at the behavioral planning level and does not describe the physical alignment of a sensor unit with spatial areas associated with the interaction context. 2.8 Industrial Cognitive Robot Arms Systems from Neura Robotics GmbH (particularly DE 10 2021 100 112 A1 and DE 10 2023 116 771 B3) are known in the field of industrial cognitive robot arms. These systems equip collaborative robot arms with multimodal sensors (cameras, force / torque sensors, audio) and machine learning. The MAiRA product has been marketed as a cognitive industrial robot since 2020. Control is pose-reactive, achieved by recognizing defined operator poses, or via image-based teach-in methods during ongoing industrial work processes. Shortcomings: These systems are geared towards task-based industrial applications (e.g., pick-and-place, welding, grinding). A state machine that autonomously switches between semantically different perceptual states without manual triggering and derives the active state from the accumulated context of an ongoing human-system interaction is not disclosed. 2.9 Summary of the shortcomings of the state of the art All known solutions share the common characteristics that they either (a) react purely to a single stimulus, (b) require manual activation by the user, (c) lack physical orientation control of the sensor unit between semantically different perceptual states, or (d) operate at the behavioral planning level without controlling the physical sensor orientation. None of the known systems derives a perceptual state from the accumulated context of an ongoing human-system interaction and, based on this, autonomously controls the physical orientation of a sensor unit to a spatial area associated with the perceptual state. 3. Object of the invention The invention is based on the objective of providing a system for environmental perception that autonomously controls the physical orientation of a sensor unit without manual triggering by a user, wherein the control is based on context information generated during ongoing human-system interaction, and wherein the system distinguishes between semantically different perceptual states, each of which is assigned different spatial areas. 4. Solution The problem is solved by a system having the features of claim 1. Advantageous embodiments are specified in claims 2 to 15. One use is specified in claim 16. The system comprises a sensor unit 10 for acquiring environmental information, at least one actuator 20 for physically changing the spatial orientation of the sensor unit 10, and a processing unit 30 with a state machine 31. The state machine 31 switches autonomously between at least two semantically distinct perceptual states Z1, Z2 without manual activation by a user. The processing unit 30 derives the currently active perceptual state from context information K generated during an ongoing human-system interaction and controls the actuator 20 such that the sensor unit 10 is aligned with a spatial area R1, R2 associated with the active perceptual state. It is essential that the state machine 31 does not perform the transition between perceptual states on the basis of a single stimulus (such as a recognized face or a speech command), but derives it from the context of the ongoing human-system interaction. In a preferred embodiment, the processing unit 30 comprises a locally executed learning-based processing module 32 that accumulates the context information K from a plurality of temporally successive interaction events and controls the state machine 31 on the basis of an interaction context derived from these accumulated events (claim 2). In a further embodiment, the processing unit 30 includes a persistent user-specific context memory 33, which stores information about user preferences and interaction histories beyond individual interaction sessions (claim 3). The state machine 31 can additionally determine the active perceptual state based on an emotional state of the user derived from acoustic and / or visual signals of the human-system interaction (claim 4). In a preferred embodiment, a physical covering device 50 is provided that physically prevents detection by the sensor unit 10 when none of the perception states are active. The physical covering device 50 can, for example, be designed as a mechanical element (such as a movable eyelid or an aperture), as an electrochromic layer (darkening glass), as a liquid crystal shutter, or as another physically acting device. The covering device 50 is controlled by the state machine 31 (claim 5). The processing unit 30 can perform all calculations required to determine the active perception state locally on the system without transferring personal data to external data processing systems (claim 6). In one embodiment, the state machine 31 includes an initiation state Z3, in which the system autonomously initiates such interaction without prior human-system interaction. The transition to the initiation state Z3 is based on context information about the presence or state of a user derived from the sensor unit 10 (claim 9). When several users are present simultaneously, the state machine 31 can determine a prioritized perceptual state by comparing user-specific context information from the persistent context memory 33 with the processing unit 30 (claim 10). The actuator 20 can execute a predetermined movement sequence of the sensor unit 10 during a transition between perceptual states, which perceptibly signals the active perceptual state to the user (claim 11). The movement sequence comprises a variable speed of the actuator 20, which is derived from the active perceptual state and the context information K (also claim 11). The state machine 31 can include a reduced perceptual state Z4 in which only part of the sensor unit 10 is active. The transition to the reduced perceptual state Z4 is derived autonomously from an inactivity phase of the human-system interaction, with the actuator 20 moving the sensor unit 10 to a predetermined rest position (claim 12). The processing unit 30 can perform an acoustic scene classification, deriving a current interaction situation from ambient noise. The state machine 31 determines the active perceptual state based on the classified interaction situation (claim 13). The learning-based processing module 32 can adjust transition probabilities between perceptual states based on success indicators derived from past human-system interactions (claim 14). The state machine 31 can additionally consider environmental conditions when determining the active perceptual state. These conditions are derived from sensor data acquired independently of the human-system interaction (claim 15). Such environmental conditions include, for example, the time of day, the lighting conditions in the room, or the ambient noise level. For example, the state machine 31 can prefer an acoustic perceptual state in darkness and a visual perceptual state in daylight. 5. Advantages of the invention The system according to the invention offers the following advantages over the prior art: Compared to reactive face tracking mechanisms and eye-tracking systems, the state machine 31 enables autonomous switching between semantically different perceptual states based on the interaction context. The system is not limited to tracking a single stimulus. Compared to voice-controlled alignment systems, the system acts proactively – the state machine 31 derives the active perceptual state from the interaction context without requiring manual triggering by the user. The initiation state Z3 also enables the system to independently initiate a human-system interaction. Compared to cloud-dependent systems, local processing enables operation without a network connection and without the transmission of personal data.The physical cover device 50 additionally offers a physically detectable data protection mechanism. The acoustic scene classification enables the determination of the perceptual state even without active verbal interaction from the user, for example, by recognizing game sounds, reading aloud, or silence. The learning state transitions enable increasingly user-adapted perceptual control by prioritizing state transitions that lead to successful interactions. The variable movement speed of the actuator 20 during state transitions allows for user-understandable signaling of the system state through the nature of the physical movement. 6. Examples of Implementation 6.1 First embodiment: Interactive children's toy (Figs. 1, 3, 4) The invention is explained below with reference to a first embodiment in which the system is designed as an interactive children's toy in the form of an owl figure (claims 1-7, 9-15). 6.1.1 System structure Fig. 1 shows a schematic representation of system 1 according to the first embodiment. System 1 comprises a housing 2 in the shape of an owl figure. A sensor unit 10, a microphone 11, a loudspeaker 12, and a mechanical cover 50 are arranged in a movable head section 3 of the housing 2. The sensor unit 10 includes a camera module 13 with a resolution of at least 640×480 pixels. The camera module 13 is located behind a transparent cover in an eye area 4 of the owl figure. An actuator 20 is designed as a servo motor (for example, of type MG90S with a torque of at least 1.8 kg·cm). The actuator 20 is connected to the head section 3 via a gearbox 21 and enables the head section 3 to rotate about a vertical axis 22 with a swivel range of at least 120 degrees. A second actuator 23 enables the head section 3 to tilt about a horizontal axis 24. A processing unit 30 is implemented as a single-board computer (for example, a Raspberry Pi Zero 2W with a 1 GHz quad-core processor and 512 MB of RAM). The processing unit 30 is located in the body section 5 of the housing 2 and is connected via signal lines to the sensor unit 10, the microphone 11, the speaker 12, and the actuator 20. The processing unit 30 comprises a state machine 31, which is executed as a software module on the processing unit 30, a learning-based processing module 32, a persistent user-specific context memory 33, which is stored on a non-volatile storage medium 34 (for example, an SD card), a speech processing module 35, and a content filter 36. The mechanical cover 50 is designed as a movable eyelid over the eye area 4 of the owl figure. The cover 50 is controlled by an actuator 51 (for example, a micro servo motor). In the closed position, the cover 50 completely conceals the camera module 13. The closing action is recognizable to the user as the owl figure's closed eyes. A power supply 60 (accumulator with charging device) supplies all components with electrical energy. 6.1.2 States of Perception Fig. 3 shows the state diagram of the state machine 31. The state machine 31 manages at least the following perception states: First perception state Z1 (user attention). The sensor unit 10 is directed towards the spatial area R1 in which the user is located. The camera module 13 detects facial features and body posture. The microphone 11 detects speech utterances. This state is active during direct verbal interaction. Second perception state Z2 (environment perception). The sensor unit 10 is directed towards a spatial area R2 that differs from the spatial area R1. The system detects the user's environment, for example, toys, books, or other people. Third perception state Z3 (initiation state according to claim 9). The sensor unit 10 is directed towards an input area R3.The system detects the presence of a user and autonomously initiates a human-system interaction, for example, by greeting the user via speaker 12. The transition to initiation state Z3 occurs without prior human-system interaction, based on contextual information K about a detected presence (for example, footsteps, a recognized voice, or a change in the ambient noise level upon entering a room). Initiation state Z3 allows the system, for example, to greet a child upon returning home or to proactively address a child doing homework, without the child having previously activated the system. The following applies: Even if the system processes a user gesture (for example, waving or pointing) as context information K, this is not a manual trigger in the sense of claim 1. The state machine 31 processes the recognized gesture as one of several context information items and decides autonomously on the state transition. Fourth perceptual state Z4 (reduced perceptual state according to claim 12). Only the microphone 11 is active; the camera module 13 is deactivated. The cover 50 conceals the camera module 13. The actuator 20 moves the headpiece 3 into a predetermined resting position (slightly tilted downwards, eyes closed). The transition to perceptual state Z4 is derived autonomously from an inactivity phase. 6.1.3 Context Information and State Transitions The state machine 31 receives context information K from the sensor unit 10 and the microphone 11 and uses this information to determine the active perceptual state. The context information K includes: Linguistic context information. The learning-based processing module 32 analyzes the content of the user's speech utterances. If the user refers to an object or direction ("Look over there"), the state machine 31 initiates a transition from Z1 to Z2. Emotional context information (claim 4). The learning-based processing module 32 derives an emotional state of the user from acoustic and / or visual signals of the human-system interaction. Acoustic features include prosody, speech rate, pitch, and volume of the speech utterances. Visual features, which are captured via the camera module 13, include facial expression, mimetic gestures (e.g., raised eyebrows, downturned corners of the mouth), gaze direction, and posture of the user.Acoustic and visual features can be used individually or multimodally to determine emotional state. Upon detecting excitement or sadness, the state machine 31 can maintain perceptual state Z1 instead of switching to Z2. Acoustic scene classification (claim 13). The processing unit 30 classifies the acoustic scene into one of several predetermined categories: • Active play (laughter, sounds of movement, increased noise level) • Quiet activity (soft talking, page turning, low volume) • Reading aloud (rhythmic speaking of an adult voice) • Silence following prior activity. The classified interaction situation is fed into state machine 31 as context information K. For example, if the scene classification detects a transition from active play to silence, state machine 31 can activate the initiation state Z3. Accumulation (Claim 2). The learning-based processing module 32 accumulates the contextual information K across a plurality of temporally successive interaction events. The state machine 31 does not react to a single event, but considers the course of the interaction. For example, the state machine 31 recognizes from a sequence of interaction events (user talks about a book → user points in a direction → user asks "Do you see that?") that a transition from Z1 to Z2 is appropriate. Persistent context information (Claim 3). The context memory 33 stores the following for each recognized user: • A user identifier (derived from voice characteristics, without storing raw biometric data) • Preferences (preferred topics, favorite stories) • Interaction history (times and duration of past interactions) • Learned state transition probabilities (Claim 14) If the system recognizes a known user, the associated preferences and learned transition probabilities are loaded. Learning state transitions (claim 14). The learning-based processing module 32 adjusts the transition probabilities between perceptual states based on success indicators. One success indicator is, for example, the duration of the subsequent user interaction after a state transition: If the transition from Z1 to Z2 is followed by sustained interaction (the user continues speaking, shows interest), the transition probability for this transition is increased in similar contexts. If the interaction is interrupted (silence, turning away), the transition probability is decreased. The adjusted transition probabilities are stored in the context memory 33. 6.1.4 Multi-user prioritization According to claim 10, the state machine 31 determines a prioritized perceptual state when multiple users are present simultaneously. The prioritization is based on user-specific context information from the context memory 33. For example, a user with whom an ongoing interaction exists is prioritized, while an arriving user is initially greeted via the loudspeaker 12 without the sensor unit 10 changing its orientation. The state machine 31 can then decide, based on the context information K, to orient the sensor unit 10 towards the second user when the latter enters the interaction. 6.1.5 Status signaling through movement According to claim 11, the actuator 20 executes a predetermined movement sequence during a transition between perception states: • Transition from Z4 (reduced) to Z3 (initiation): The head 3 rises from its rest position, and the cover 50 exposes the camera module 13. • Transition from Z1 (user facing) to Z2 (environment detection): The head 3 rotates away from the user in the direction of the spatial area R2 to be detected. • Transition to Z4 (reduced): The head 3 lowers, and the cover 50 closes. Also according to claim 11, the movement sequence comprises a variable speed of the actuator 20, which is derived from the active perceptual state and the contextual information K. If user excitement is detected (rapid speech, increased volume), the movement in Z1 is fast. If calmness is detected (soft speech, even prosody), the movement is slow. During a transition to Z4, the movement is always slow to signal the figure gently falling asleep. 6.1.6 Processing without external data transfer According to claim 6, the processing unit 30 performs all calculations necessary to determine the active perceptual state without transferring personal data to external data processing systems. In a first embodiment, the learning-based processing module 32 is stored on and executed on the processing unit 30; in this embodiment, the system 1 is functional without a network connection. In a second embodiment, the processing unit 30 can be connected to a separate computing unit via a non-public local network (for example, via a home wireless network, a dedicated wireless protocol, or a wired connection within the same residential unit). In this embodiment, some of the calculations, in particular computationally intensive inference steps of the learning-based processing module 32, can be performed on the separate computing unit, while time-critical control functions remain on the processing unit 30. In this second embodiment as well, no personal data is transferred to data processing systems outside this local network. 6.1.7 Age-appropriate language processing and content filtering According to claim 7, the processing unit 30 comprises a speech processing module 35 designed for interaction with children. The speech processing module 35 uses age-appropriate vocabulary and sentence structure. A content filter 36 checks the output generated by the system for topics that are not age-appropriate and suppresses them. 6.1.8 Example of an interaction process Fig. 4 shows the time sequence of an example interaction. Step 1: System 1 is in the reduced perception state Z4. The headpiece 3 is in the resting position, the cover 50 is closed. Only the microphone 11 is active. Step 2: The acoustic scene classification detects sounds indicating the presence of a person (footsteps, voice). The state machine 31 transitions to the initiation state Z3. Step 3: The cover 50 opens. The actuator 20 slowly lifts the head section 3 and rotates it towards the detected noise source (room area R3). The speed of movement is moderate (signaling waking up). Step 4: The camera module 13 recognizes a user's face. The processing unit 30 compares the recognized features with the context memory 33 and identifies the user. The system then plays a personalized greeting through the speaker 12. Step 5: The state machine 31 switches to Z1 (user orientation). The actuator 20 aligns the head section 3 towards the user. The human-system interaction begins. Step 6: During the interaction, the learning-based processing module 32 accumulates context information K. The user talks about a book and points in a direction. The state machine 31 derives a transition to Z2 from the accumulated context information. Step 7: The actuator 20 quickly rotates the head section 3 in the indicated direction (room area R2). The increased speed of movement signals interest. The sensor unit 10 detects the object. The system describes the detected object via the speaker 12. Step 8: The state machine 31 switches back to Z1. The actuator 20 rotates the head section 3 back towards the user. Step 9: After an inactivity period of, for example, 5 minutes, the state machine 31 switches to Z4. The actuator 20 slowly moves the head 3 to its rest position. The cover 50 closes. The context memory 33 stores the interaction information and the updated transition probabilities. 6.2 Second embodiment: Stationary personal assistant (Fig. 2) The invention is explained below with reference to a second embodiment in which the system is designed as a stationary personal assistant (claims 1-6, 8, 9, 11, 13-15). 6.2.1 System structure Fig. 2 shows a schematic representation of system 100. System 100 comprises a housing 102 designed for desktop placement. A movable upper section 103 of the housing 102 carries a sensor unit 110 with a camera module 113 and a microphone 111. An actuator 120 enables the upper section 103 to pivot about a vertical axis 122. A second actuator 123 enables it to tilt about a horizontal axis 124. A processing unit 130 is arranged in the housing 102 and comprises a state machine 131, a learning-based processing module 132 and a persistent context memory 133. 6.2.2 States of Perception The state machine 131 uses two of the perceptual states Z1 and Z2 shown in Fig. 3, which are specified below for the application context of a personal assistant. Perception state Z1 (in the second embodiment: user state). The sensor unit 110 is directed towards the spatial area R101 in which the user is located (chair in front of the desk). The system 100 detects the user's facial expression, head posture and gaze direction. Perception state Z2 (in the second embodiment: work material). The sensor unit 110 is directed towards the spatial area R102, which comprises the desk surface in front of the user. The system 100 detects objects or environmental elements in the user's immediate vicinity, for example, documents, screen content, cooking utensils, or medication. 6.2.3 Functional sequence During an interaction, the system 100 is in perceptual state Z1. The learning-based processing module 132 accumulates contextual information from the user's speech. If the user places a document on the desk and comments on its content, the state machine 131 derives a transition to Z2 from the accumulated contextual information. The actuator 120 tilts the upper part 103 downwards onto the desk surface. The sensor unit 110 detects the document. The processing unit 130 can perform text recognition and play back the recognized content via a speaker 112. After document capture is complete, the state machine 131 returns to Z1. The actuator 120 realigns the upper part 103 towards the user again. 7. List of reference symbols 1 System (children's toy) 2 Housing (owl figure) 3 Head (movable) 4 Eye area 5 Body 10 Sensor unit 11 Microphone 12 Speaker 13 Camera module 20 Actuator (servo motor, pan movement) 21 Gearbox 22 Vertical axis 23 Second actuator (tilt movement) 24 Horizontal axis 30 Processing unit (single-board computer) 31 State machine (software module) 32 Learning-based processing module 33 Persistent user-specific context memory 34 Non-volatile storage medium (SD card) 35 Speech processing module 36 Content filter 50 Mechanical cover device (eyelid) 51 Cover device actuator 60 Power supply Z1 First perceptual state (user attention) Z2 Second perceptual state (environment perception) Z3 Third perceptual state (initiation state) Z4 Fourth perceptual state (reduced state) R1 First Room area (position of the user) R2 Second room area (surroundings) R3 Third room area (entrance area) B1 Transition condition Z4 →Z3 (acoustic presence detection) B2 Transition condition Z3 → Z1 (user identification) B3 Transition condition Z1 → Z2 (reference to object in interaction context) B4 Transition condition Z2 → Z1 (detection of the referenced object) B5 Transition condition Z1 → Z4 (inactivity of human-system interaction) B6 Transition condition Z3 → Z4 (timeout without user identification) 100 System (stationary personal assistant) 102 Housing 103 Top part (movable) 110 Sensor unit 111 Microphone 112 Speaker 113 Camera module 120 Actuator (swivel movement) 122 Vertical axis 123 Second actuator (tilt movement) 124 Horizontal axis 130 Processing unit 131 State machine 132 Learning-based processing module 133 Persistent context memory R101 First room area (user position) R102 Second room area (desk surface) _Note: The perceptual states are generally represented in Fig. 3 with the reference symbols Z1 to Z4. The second embodiment uses states Z1 (here: user state) and Z2 (here: working material)._ 8. Character description Fig. 1: Schematic representation of system 1 as an interactive children's toy (owl figure) with sensor unit 10, actuator 20, processing unit 30, and mechanical cover device 50. Fig. 2: Schematic representation of system 100 as a stationary personal assistant with sensor unit 110 and actuators 120, 123. Fig. 3: State diagram of the state machine 31 with the perception states Z1, Z2, Z3, Z4 and the transition conditions B1 (Z4 → Z3), B2 (Z3 → Z1), B3 (Z1 → Z2), B4 (Z2 → Z1), B5 (Z1 → Z4), and B6 (Z3 → Z4, shown as a dashed arrow; timeout transition). The double frame designates Z1 as the main state. The meaning of transition conditions B1 to B6 is given in the reference numeral list and is explained in sections 6.1.2 and 6.1.3. Fig. 4: Time sequence of an exemplary interaction according to the first embodiment (steps 1-9) showing the perception state transitions and actuator movements. QUOTES INCLUDED IN THE DESCRIPTION This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature US 11,148,296 B2

[0006] US 12,136,431 B2

[0007] WO 2013 / 178741

[0008] WO 2015 / 158876

[0008] EP 3 894 972 A1

[0010] EP 3 691 766 B1

[0012] US 12,197,409 B2

[0014] EP 4 143 675 A1

[0017] EP 4 578 604 A1 [0020, 0021]DE 10 2021 100 112 A1

[0022] DE 10 2023 116 771 B3

[0022]

Claims

- Main claim (product claim) System for context-sensitive environmental perception, comprising a sensor unit for capturing environmental information, at least one actuator for physically changing the spatial orientation of the sensor unit, and a processing unit with a state machine, characterized in that the state machine switches autonomously between at least two semantically different perception states without manual triggering by a user, wherein the processing unit derives the currently active perception state from context information generated during an ongoing human-system interaction and controls the actuator in such a way that the sensor unit is aligned with a spatial area assigned to the active perception state. - Subclaim: accumulating learning-based processing module system according to claim 1, characterized in that the processing unit comprises a locally executed learning-based processing module that accumulates the context information from a plurality of temporally successive interaction events and controls the state machine on the basis of an interaction context derived from these accumulated events. - Subclaim: persistent user memory system according to claim 1 or 2, characterized in that the processing unit comprises a persistent user-specific context memory which stores user-specific information, in particular user preferences and interaction histories, beyond individual interaction sessions and takes these into account when determining the active perceptual state. - Subclaim: emotional state recognition system according to one of claims 1 to 3, characterized in that the state machine additionally determines the active perceptual state on the basis of an emotional state of the user derived from acoustic and / or visual signals of the human-system interaction. - Subclaim: data protection controlled covering device system according to one of claims 1 to 4, characterized in that a physical covering device is provided which physically prevents detection by the sensor unit when none of the perception states is active, and which is controlled by the state machine. Subclaim: Processing without external data transfer system according to one of claims 1 to 5, characterized in that the processing unit performs all calculations required to determine the active perception state without transferring personal data to external data processing systems. - Subclaim: Design as a children's toy system according to one of claims 1 to 6, characterized in that the system is designed as an interactive children's toy with age-appropriate speech processing and / or a content filter for non-age-appropriate topics. - Subclaim: Design as a personal assistant system according to one of claims 1 to 6, characterized in that the system is designed as a personal assistant, the first perception state of which is directed towards the detection of a user state and the second perception state of which is directed towards the detection of objects or environmental elements in the environment of the user. - Subclaim: proactive interaction initiation system according to one of claims 1 to 8, characterized in that the state machine comprises an initiation state in which the system autonomously initiates such interaction without prior human-system interaction, wherein the transition to the initiation state is based on context information about the presence or state of a user derived from at least one sensor of the system and the actuator aligns the sensor unit with the detected user. - Subclaim: Multi-user prioritization system according to claim 3, characterized in that the state machine determines a prioritized perceptual state when several users are present simultaneously, by comparing user-specific context information from the persistent context memory and controlling the actuator in such a way that the sensor unit is aligned with the prioritized user. - Subclaim: State signaling by movement system according to one of claims 1 to 10, characterized in that the actuator, during a transition between perceptual states, executes a movement sequence of the sensor unit with a variable speed derived from the active perceptual state and the context information, which perceptibly signals the active perceptual state to the user. - Subclaim: reduced perception state system according to one of claims 1 to 11, characterized in that the state machine comprises a reduced perception state in which only a part of the sensor unit is active, wherein the transition to the reduced perception state is derived autonomously from an inactivity phase of the human-system interaction and the actuator moves the sensor unit to a predetermined rest position. - Subclaim: acoustic scene classification system according to one of claims 1 to 12, characterized in that the processing unit performs an acoustic scene classification which derives a current interaction situation from ambient noise, and the state machine determines the active perceptual state on the basis of the classified interaction situation. - Subclaim: learning state transition system according to claim 2, characterized in that the learning-based processing module adapts transition probabilities between the perceptual states on the basis of success indicators derived from past human-system interactions. - Subclaim: Environmental conditions as a state influencing factor System according to one of claims 1 to 14, characterized in that the state machine additionally takes into account environmental conditions derived from sensor data which are acquired independently of the human-system interaction when determining the active perception state. - Use claim: Use of a system according to one of claims 1 to 15 in an interactive device, in particular as a children's toy, personal assistant, smart home device, care assistance system or learning support system, for autonomous context-controlled alignment of the sensor unit based on context information generated during a human-system interaction.