Device, system, computer-implemented method and computer program for monitoring a specific spatial and / or functional area by means of a data interface
A system combining image and audio data using pre-trained models for environmental monitoring addresses low detection accuracy and adaptability issues, enhancing incident detection and response through linked textual descriptions.
Patent Information
- Application Number
- DE102024201038
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-06
- Publication Date
- 2025-08-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing machine monitoring systems for environments struggle with low accuracy in detecting incidents, leading to missed detections and false alarms, and are unable to adapt to new types of incidents that were not previously considered.
A system combining image and audio data using pre-trained machine learning models to generate linked textual descriptions, enabling improved incident detection and response through a hardware unit that integrates a basic model to logically link image and audio embeddings, allowing for real-time monitoring and action determination.
Enhances the accuracy and adaptability of environmental monitoring by integrating image and audio data, facilitating timely and appropriate responses to incidents, thereby improving safety and efficiency.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a device, a system, a computer-implemented method and a computer program for monitoring a specific spatial and / or functional area by means of a data interface.
[0002] The following definitions, descriptions and statements retain their respective meaning for and apply to the entire disclosed subject matter of the invention.
[0003] In the field of computer vision, image classification, machine learning models are known that can detect, classify, localize and / or track objects in input images, for example camera images.
[0004] State-of-the-art machine learning models such as BERT, DALL-E, and GPT series are emerging. These models are trained on large data sets and can be adapted to a variety of downstream tasks. Such models are called foundation models, see, for example, ar-Xiv:2108.07258v3 [cs.LG] 12 Jul 2022. One example of a foundation model is language models, also called large language models. A language model is a machine learning model that is trained in a data-driven training procedure to model a sequence of elements of a sequence, for example, letters or words in natural language texts, see, for example, ar-Xiv:1706.03762v7 [cs.CL] 2 Aug 2023. A publicly available language model is, for example, LLaMA, a collection of foundation language models from Meta AI, see, for example, arXiv:2302.13971v1 [cs.CL] 27 Feb 2023.Using prompt engineering, language models can process natural language and demonstrate logical relationships. For example, a task, such as a question, is presented to the language model in text form via an interface, such as an input field.
[0005] In arXiv:1706.10006v2 [cs.SD] 24 Oct 2017, a computer-implemented method for the automatic text-based description of audio signals, also called automated audio captioning, abbreviated ACC, see also https: / / dcase.community / challenge2023 / task-automated-audio-captioning-and-language-based-audioretrieval, is disclosed. This automatically generates a textual description for an audio signal, e.g., obtained from a microphone, that is as close as possible to a human-assigned audio signal. This is achieved by an artificial neural network comprising an encoder-decoder structure. Between the encoder and decoder, the artificial neural network includes an intermediate layer that receives as input an output from the encoder at a time i and an output from the decoder at the previous time i-1, and generates as output a prediction for the decoder input at time i.This intermediate layer is called a soft alignment model or soft attention mechanism and helps the decoder focus on different encoder outputs for calculating the output at time i, taking into account the decoder's previous states and predictions at time i-1. This artificial neural network can be trained on a dataset such as that disclosed in ar-Xiv:1910.09387v1 [cs.SD] 21 Oct 2019.
[0006] Image and / or audio recognition using machine learning models as described above enables environmental monitoring. However, to date, the accuracy of detecting incidents—for example, accidents, robberies, burglaries, vandalism, the arrival of visitors, speeding, failure to yield, weather events, or trams—by humans is significantly better than that of machine-based surveillance. Therefore, many incidents continue to go undetected, or many false alarms occur. Subsequent human analysis of image / video material and / or audio recordings is usually sufficient to detect the incident. This means that sensor technology, including camera and / or acoustic sensors, is often good enough, but computer-implemented analysis methods are too poor to detect all types of incidents.The above list of incidents that could be detected using image / video and / or audio material is not exhaustive, as such a list is impossible to be complete, as new, previously unknown incidents may occur over time, for example a still fictitious event such as a hovercar falling below its flight altitude.
[0007] Image / video and acoustic data can be combined in several ways. The terms "acoustic data" and "audio data" are used interchangeably. The fusion of the two data streams plays a key role. This fusion can occur early, also called early fusion, where all raw data is input into an artificial neural network, for example, or late fusion, where an artificial neural network is provided for each image sensor and acoustic sensor, which classifies the respective data, and a fusion only occurs after the respective classification. However, early fusion requires a large amount of test data, which would have to be provided for training.
[0008] The object of the invention was to find out how machine monitoring of an environment can be improved, in particular to increase safety, increase the efficiency of a monitoring system and / or minimize error detection in the monitoring system.
[0009] The subject matter of the independent and subordinate claims each solves this problem. Advantageous embodiments of the invention emerge from the definitions, the subclaims, the drawings, and the description of preferred embodiments.
[0010] According to one aspect, the invention provides a device for monitoring a specific spatial and / or functional area via a data interface. The area can also be called the surroundings or environment. The area can be characterized by: • Spatial boundaries: The environment is a clearly defined physical space determined either by natural boundaries, such as an ecosystem or a geographical location, or by artificial boundaries, such as the interior of a building, the boundary of a traffic space, for example a roadway, or an industrial site, or the defined areas of a facility. • Functional properties: The environment is characterized by its functional use or purpose, such as industrial production, traffic monitoring, health monitoring, agricultural monitoring, security monitoring or environmental monitoring. • Interacting elements: The environment includes all relevant objects that interact or are present within the defined space. These include, for example, road users, vehicles, people, environmental factors, and / or machines. • Dynamic conditions: The environment may be characterized by dynamic or changing conditions, such as variable environmental conditions, changing operating conditions or human activities. • Measurement and monitoring points: The environment may include specific points or areas where measurements or monitoring are carried out to collect data relevant to the monitoring function.
[0011] Monitoring such areas is a concrete technical problem.
[0012] The claimed objects can be used in a wide range of areas, including environmental protection, industrial process control, building safety, transportation systems, automated or autonomous driving, agriculture, and healthcare. They play a crucial role in ensuring safety, efficiency, and sustainability in these fields.
[0013] The device comprises a hardware unit executing a pre-trained baseline model. The hardware unit, including the first hardware unit and the second hardware unit, may be an integrated circuit or an application-specific integrated circuit comprising one or more CPUs, GPUs, FPGAs, or other computing and / or memory components. The hardware unit is a technical means for performing the monitoring. The baseline model is, in one aspect, a language model. For example, the language model is the publicly available and usable language model LLaMA by Meta Al, which was trained, for example, on datasets obtained from web crawling. Data that can be used for pre-training are disclosed, for example, in arXiv:2302.13971v1 [cs.CL] 27 Feb 2023.According to one aspect, the basic model can also include a vision language model, which directly understands video and image data at the pixel level and does not require text comprehension. For example, a further development of LLaMA, namely LLaMA-VID, a vision language model based on the LLaMA architecture, is disclosed in arXiv:2311.17043v1 [cs.CV] 28 Nov 2023. According to one aspect, the basic model can also include a speech language model, which directly understands audio data and does not require text comprehension. For example, arXiv:2308.16692v1 [cs.CL] 31 Aug 2023 discloses a Unified Speech Language Model. According to one aspect, the data interface is a speech interface. The basic model receives a first text and a second text via the speech interface. In this context, the data interface corresponds to a text-based data interface.
[0014] According to a further aspect, an encoder of the first machine learning model encodes the image data as vectors into a vector space, thereby obtaining image embeddings. The vector space is also called an embedding space or latent feature space. A decoder of the first machine learning model decodes the vectors into a first text. The first hardware unit outputs the first text as the first character sequence. The hardware unit receives the first text as the first character sequence via the data interface. An encoder of the second machine learning model encodes the audio data as vectors into a vector space, thereby obtaining audio embeddings. The vector space is also called an embedding space or latent feature space. The encoder of the first machine learning model and the second machine learning model, as well as the decoder of the first machine learning model and the second machine learning model, can share a common vector space, a so-called shared embedding space.A decoder of the second machine learning model decodes the vectors into a second text. The second hardware unit outputs the second text as the second character sequence. The hardware unit receives the second text as the second character sequence via the data interface.
[0015] According to a further aspect, the encoder of the first machine learning model encodes the image data as vectors in a vector space and thus obtains image embeddings. The first hardware unit outputs the image embeddings as the first character sequence. The hardware unit receives the image embeddings as the first character sequence via the data interface. The encoder of the second machine learning model encodes the audio data as vectors in a vector space and thus obtains audio embeddings. The second hardware unit outputs the audio embeddings as the second character sequence. The hardware unit receives the audio embeddings as the second character sequence via the data interface.
[0016] In the two preceding embodiments, the data interface is a token-based interface. This represents an alternative to the text-based interface between the camera and audio raw data and the base model. This allows the input token embeddings to be used as a data interface. For this purpose, the first machine learning model can be trained to bring comparable image data in the same embedding space into proximity with existing text embeddings of the base model. Furthermore, the second machine learning model can be trained to bring comparable audio data in the same embedding space into proximity with existing text embeddings of the base model. In one embodiment, only the encoder of the first and second machine learning models is trained, which does not change the performance of the base model. Following this, a prompt according to one aspect looks like this: "You are an automatic surveillance system. Audio and camera signals are available to monitor your surroundings. Your camera image shows: <Spezialtoken, die den Input-Text-Encoder überspringen und direkt die embeddings aus dem Kamerabild übernehmen> Listen to your audio data: <Spezialtoken, die den Input-Text-Encoder überspringen und direkt die embeddings aus dem Audiosignal übernehmen> . Your speech recognition recognizes: <Spezialtoken, die den Input-Text-Encoder überspringen und direkt die embeddings aus der Spracherkennung übernehmen> Your possible actions are: Do nothing / Notify the owner / Raise the alarm / Initiate diversionary measures such as simulating presence (turning on lights in the house, playing sounds over loudspeakers) / Inform the police.
[0017] According to a further aspect, particularly when the base model comprises a vision language model and / or a speech language model as described above, two different tokens are read in per frame for image data via the token-based interface: a context token that encodes the overall image and a content token that encodes visual features per frame. Two tokens in base models are disclosed, for example, in arXiv:2311.17043v1 [cs.CV] 28 Nov 2023. Analogously, two different tokens can be read in for audio data, namely semantic tokens and acoustic tokens, as disclosed, for example, in arXiv:2308.16692v1 [cs.CL] 31 Aug 2023.
[0018] The first character sequence is received by a first hardware unit that reads image data from an image sensor and outputs the first character sequence. The image sensor can be, for example, a camera sensor, ultrasonic sensor, lidar sensor, or radar sensor. The first hardware unit executes a first machine learning model that has been trained on image data to detect, classify, localize, and / or track objects and converts them into a correspondingly descriptive character sequence. The first character sequence can be text, for example, a natural language text. This is also called image-to-text conversion. The first character sequence can also be an image embedding, as described above. The first machine learning model can be an artificial neural network. The training can be supervised learning. Tracking is also called tracking.Classification, localization and / or tracking can be achieved using convolutional layers, recurrent layers or transformer architecture.
[0019] The second character sequence is obtained from a second hardware unit that reads audio data from an acoustic sensor and outputs the second character sequence. The acoustic sensor can be, for example, a microphone. The second hardware unit executes a second machine learning model that has been trained on audio data to detect, classify, localize, and / or track events and convert them into a correspondingly descriptive character sequence. The second character sequence can be a text, for example, a natural language text, or an audio embedding, as described above. The second machine learning model is, for example, an artificial neural network and, in one aspect, comprises an encoder-decoder structure as described in arXiv:1706.10006v2 [cs.SD] 24 Oct 2017 under section 2 and in particular in Fig. 1. By this explicit reference, the disclosure of arXiv:1706.10006v2 is incorporated into the present disclosure. For example, the second machine learning model was trained on the Clotho dataset, as disclosed in arXiv:1910.09387v1 [cs.SD] 21 Oct 2019.
[0020] The first hardware unit and the second hardware unit can each be spatially separated from the hardware unit. According to one aspect, the first hardware unit with the image sensor and the second hardware unit with the acoustic sensor each form an integrated unit, which are arranged in the spatial and / or functional area to be monitored or are each arranged such that their respective detection range covers the spatial and / or functional area or parts thereof. The hardware unit can then be arranged remotely. The data exchange between the first hardware unit and the second hardware unit, each with the hardware unit, can be wired or carried out using wireless technology.
[0021] The image sensor and acoustic sensor are technical devices that monitor physical or chemical parameters and provide real-time data about the area. By analyzing the image and audio data with the first and second machine learning models, measurement results are technically evaluated and technically relevant conclusions for the monitoring are drawn. These conclusions are converted into speech according to one embodiment. For example, an image recording may show a man dressed in black climbing over a fence to enter a property. The first machine learning model reads this image and outputs the first text: "Your camera image shows a man dressed in black climbing over a fence." For example, an audio recording from this area may include mumbling and / or metal-on-metal clanking and a speech recognition of "Quick, over here."The second machine learning model reads these sounds and outputs them in the form of a second text: "Your audio data hears mumbling, metal-on-metal clanking, and 'Quick, over here'."
[0022] Based on the pre-training, the baseline model logically links the first character sequence, for example, the first text, and the second character sequence, for example, the second text. For example, the pre-trained baseline model understands the logical connection that a burglar intends to enter a property from the text inputs "Your camera image shows a man dressed in black climbing over a fence" and "Your audio data hears mumbling, metal-on-metal clanking, and 'Quick, over here.'" The camera recordings, together with the corresponding audio recordings, make a burglary situation seem likely. The text inputs thus enable the use of a baseline model, for example, a language model, for surveillance.The advantage of the baseline model is that it enables machine monitoring and that the quality of monitoring is improved by exploiting the trained logical connections of the baseline model. Since a baseline model does not need to be trained on specific events, such as the detection of a break-in, theft, or vandalism of a vehicle, a baseline model can cover a wide variety of application scenarios, especially those events that were not considered at the time the baseline model was created.By logically linking the first character sequence, for example, the first text, and the second character sequence, for example, the second text, in the basic model, the model is specifically developed. The data interface, implemented as a token-based interface in the case of image and audio embeddings, which reads the image embeddings and the audio embeddings, or as a voice interface that reads the first text and the second text, enables a multimodal connection of several sensor technologies, such as image technology and acoustic technology, to a basic model. Through the special evaluation of the sensor data, for example, in text form, in real time, the basic model becomes visible and determined by technical conditions outside the basic model.One advantage of the voice interface, i.e. the textual description of the image data and the audio data, is the improved configuration of the surveillance that it achieves.
[0023] Based on the link, the basic model determines an action to ensure the safety of people and objects in the monitored area. Thus, the surveillance fulfills a specific technical purpose, namely improving the security of people and objects. This purpose, i.e., determining the action, goes beyond mere data processing. In the case of the burglary scenario described above, the action could be, for example, "trigger alarm" or "inform police."
[0024] The device also includes a communication interface through which the action is transmitted to a user or to a control device. The communication interface enables the device to interact with the user or the control device in real time. If the user receives the action "Inform police", they will call the police. The device thus provides the user with information that helps the user in the decision-making process, take action, or control processes. If the control device receives the action "Trigger alarm", where the control device is, for example, a control device for smart home applications, the control device can, for example, turn on lights in a house or lock the property. This uses control mechanisms that respond to the monitored data to regulate or adapt technical processes.
[0025] According to another aspect, actions are stored in the basic model, and based on the link, the basic model selects the appropriate action from the stored actions. For example, the following actions are stored, for example, in a digital storage device: "Do nothing / Send notification to owner / Trigger alarm / Initiate diversionary measures such as a presence simulation (turn on lights in the house, play sounds over loudspeakers) / Inform the police." This allows the basic model to determine an action particularly quickly.
[0026] If no suitable action is available, the basic model can, according to one aspect, select the next best action and add a NEWACTION to the user in an output, prompting the user to describe which new action should be added for future events. This means that if no suitable action can be selected from the stored actions, the basic model determines a new action for this case or prompts the user via text output to determine a new action, and the new action is stored. This allows the basic model to be continuously improved for specific scenarios through user interaction with regard to determining an action.
[0027] According to a further aspect, the hardware unit is integrated into a vehicle's on-board communications network, and the communications interface transmits the action to a vehicle control unit for controlling and / or regulating a longitudinal and / or transverse movement of the vehicle, and the vehicle control unit executes the action. The vehicle can be a passenger vehicle, for example a car, shuttle, bus, or commercial vehicle, for example a truck. The on-board communications network is responsible for the flow of information between components of the vehicle, for example sensors and actuators, and the control units, for example ECUs, domain ECUs, or zone ECUs. For example, the image sensor and the acoustic sensor are vehicle sensors, and the image sensor, the first hardware unit, the acoustic sensor, the second hardware unit, and the hardware unit communicate with each other via the on-board communications network.The hardware unit transmits the action, for example, in the form of control signals or power levels to a vehicle control unit, which regulates and / or controls actuators for longitudinal and / or lateral movement. This allows the device to be used as a driver assistance system in an autonomous or non-autonomous vehicle. For example, such a device could warn the driver if an incident occurs in the immediate vicinity that the driver was not aware of, such as a cyclist falling. In another aspect, the on-board communication network is also responsible for the power supply. The on-board communication network can be a 48-volt network. The on-board network can also be a 400-volt or 800-volt network, which is advantageous for electric vehicles in terms of charging times and weight optimization.
[0028] According to a further aspect, the invention provides a system for monitoring a specific spatial and / or functional area using a data interface. The data interface can be a voice interface, for example, a text-based interface, or a token-based interface as described above in connection with the device.
[0029] The system comprises at least one image sensor and a first hardware unit. The first hardware unit executes a first machine learning model that has been trained on image data to recognize, classify, localize, and / or track objects and convert them into a correspondingly descriptive character sequence, for example, text. The first hardware unit reads the image data from the image sensor and outputs a first character sequence, for example, a first text. The first hardware unit, the image sensor, and the first machine learning model can be a first hardware unit, an image sensor, and a first machine learning model, as described above in connection with the device.
[0030] The system further comprises at least one acoustic sensor and a second hardware unit. The second hardware unit executes a second machine learning model trained on audio data to detect, classify, localize, and / or track acoustic events and convert them into a correspondingly descriptive character sequence, for example, text. The second hardware unit reads the audio data from the acoustic sensor and outputs a second character sequence, for example, a second text. The second hardware unit, the acoustic sensor, and the second machine learning model can be a second hardware unit, an acoustic sensor, and a second machine learning model as described above in connection with the device.
[0031] The system further comprises a third hardware unit executing a pre-trained basic model that reads the first character sequence, for example the first text, and the second character sequence, for example the second text, via the data interface executing a voice interface. Based on the pre-training, the basic model logically links the first character sequence, for example the first text, and the second character sequence, for example the second text. Based on the link, the basic model determines an action to ensure the safety of people and objects in the monitored area. The third hardware unit and the basic model can be a third hardware unit and a basic model as described above in connection with the device.
[0032] The system also includes a communication interface via which the action is transmitted to a user or a control unit. The communication interface can be a communication interface as described above in connection with the device.
[0033] At least in this way, the machine monitoring disclosed and implemented using the basic model does not function in isolation, but rather as an integral part of the more comprehensive system as described above. This demonstrates that the monitoring function is closely linked to the functionality and operation of the system. The application of the basic model improves system performance, the response to a detected incident, and the reliability of the system. Data acquisition using an image sensor and an acoustic sensor, analysis and further processing of the data using the first hardware unit and the second hardware unit, decision-making using the third hardware unit, and real-time communication using the communication interface are thus integrated into a single system.
[0034] The system may be deployed or used as a driving system of a vehicle, a building management system, an industrial control system, or an environmental monitoring system, or as part of any of the aforementioned systems. The vehicle may be a vehicle as described above.
[0035] According to a further aspect, the invention provides a computer-implemented method for monitoring a specific spatial and / or functional area using a data interface. The method comprises the steps: • with an image sensor obtaining an image capture from the area; • with an acoustic sensor obtaining an audio recording from the area; • Inputting the image recording into a first machine learning model that has been trained on image data to detect, classify, localize and / or track objects and converting it into a correspondingly descriptive character sequence, for example text, and outputting a first character sequence, for example a first text, that describes the area accordingly; • Inputting the audio recording into a second machine learning model trained on audio data to detect, classify, localize and / or track acoustic events and converting them into a correspondingly descriptive character sequence, for example text, and outputting a second character sequence, for example a second text, correspondingly descriptive of the area; • Inputting the first character sequence, for example the first text, and the second character sequence, for example the second text, into a hardware unit by means of the data interface, which executes a pre-trained basic model which, based on the pre-training, logically links the first character sequence, for example the first text, and the second character sequence, for example the second text, and determines an action based on the link to ensure the safety of persons and objects in the monitored area; • with a communication interface transmitting the action to a user or to a control unit.
[0036] The data interface, for example the text-based voice interface or the token-based interface, the image sensor, the acoustic sensor, the first machine learning model, the second machine learning model, the hardware unit, the basic model and / or the communication interface can be the respective components as described above in connection with the device.
[0037] According to a further aspect, the invention provides a computer program for monitoring a specific spatial and / or functional area by means of a data interface. The computer program comprises program instructions that cause a hardware unit to perform the steps of the method described above when the hardware unit executes or loads the computer program. The data interface and / or the hardware unit can be a language interface and / or a hardware unit as described above in connection with the device. The instructions of the computer program can comprise machine instructions, source code, or object code written in assembly language, an object-oriented programming language, for example, C++, or in a procedural programming language, for example, C.According to one aspect of the invention, the computer program is a hardware-independent application program that is provided, for example, via a data carrier or a data carrier signal using software over the air technology.
[0038] The invention is illustrated in the following exemplary embodiments. They show: Fig. 1 an embodiment of a device and a system disclosed here and Fig. 2 an embodiment of a method disclosed here.
[0039] In the figures, identical reference symbols designate identical or functionally similar reference parts. For clarity, only the relevant reference parts are highlighted in the individual figures.
[0040] Fig. 1 shows the device 100. Via the data interface D, the hardware unit 30 of the device 100 receives a first character sequence 11, for example a first text 11, and a second character sequence 21, for example a second text 21. The first text 11 describes an image 41. The second text 21 describes an audio file 51. The hardware unit 30 executes a pre-trained basic model 31. The basic model 31 logically links the first text 11 and the second text 21 and derives an action A therefrom. The action A is provided to a user or a control unit via the communication interface K.
[0041] The first text 11 is generated by a first hardware unit 10. The first hardware unit 10 reads an image file 41 from an image sensor 40. The first hardware unit 10 executes a first trained machine learning model 12. The first machine learning model 12 processes the image file 41 and generates the first text 11 from it. The first text 11 describes the image 41 in the respective situational context together with the associated image elements.
[0042] The second text 21 is generated by a second hardware unit 20. The second hardware unit 20 reads an audio file 51 from an acoustic sensor 50. The second hardware unit 20 executes a second trained machine learning model 22. The second machine learning model 22 processes the audio file 51 and generates the second text 21 from it. The second text 21 describes the audio file 51 semantically and acoustically.
[0043] The image sensor 40 and the acoustic sensor 50 sense an area. This sensing allows the area to be monitored.
[0044] The system 200 integrates the device 100, the image sensor 40, the first hardware unit 10, the acoustic sensor 50, and the second hardware unit 20.
[0045] The system 200 can be deployed or used as a driving system of a vehicle, a building management system, an industrial control system, an environmental monitoring system, or as a part of any of the aforementioned systems, and in each case implements machine monitoring.
[0046] Fig.2 schematically shows the method for monitoring a specific spatial and / or functional area using the data interface D. In a step V1, image recordings from the area are obtained using the image sensor 40. In a step V2, audio recordings from the area are obtained using the acoustic sensor 50. In a step V3, the image recordings are input into the first machine learning model 12, and a first text 11 describing the area according to the image recording is output. In a step V4, the audio recordings are input into the second machine learning model 22, and a second text 21 describing the area according to the audio recording is output. In a step V5, the first text 11 and the second text 21 are input into the hardware unit 30 using the data interface D into the pre-trained basic model 31. The basic model 31 logically links the first text 11 and the second text 21.Based on the link, an action A is performed to ensure the safety of people and objects in the monitored area. In step V6, the action A is transmitted to a user or to a control device via the communication interface K. The method can be performed using the device 100 or the system 200. Reference symbol 100 device 200 systems D data interface 10 first hardware unit 11 first character sequence 12 first machine learning model 20 second hardware unit 21 second character sequence 22 second machine learning model 40 image sensor 41 image data 50 acoustic sensor 51 audio data 30 hardware unit 31 Basic model A Action K Communication interface V1-V6 process steps QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature
[0000] Xiv:1706.03762v7 [cs.CL] Aug 2, 2023
[0004] arXiv:2302.13971v1 [cs.CL] 27 Feb 2023 [0004, 0013] arXiv:1706.10006v2 [cs.SD] 24 Oct 2017 [0005, 0019] arXiv:2311.17043v1 [cs.CV] 28 Nov 2023
[0017]
Claims
[1] Device (100) for monitoring a specific spatial and / or functional area by means of a data interface (D), the device (100) comprising: • a hardware unit (30) executing a pre-trained base model (31) which comprises a first character sequence (11), o the first character sequence (11) is obtained from a first hardware unit (10) which reads in image data (41) from an image sensor (40) and outputs the first character sequence (11), wherein the first hardware unit (10) executes a first machine learning model (12) which has been trained on image data to recognize, classify, localize and / or track objects and to convert them into a correspondingly descriptive character sequence, and a second character sequence (21), o the second character sequence (21) is obtained from a second hardware unit (20) which reads in audio data (51) from an acoustic sensor (50) and outputs the second character sequence (21), wherein the second hardware unit (20) executes a second machine learning model (22) which has been trained on audio data to recognize, classify, localize and / or track events and to convert them into a correspondingly descriptive character sequence, receives via the data interface (D), based on the pre-training, the first character sequence (11) and the second character sequence (21) are logically linked and, based on the link, determines an action (A) to ensure the safety of persons and objects in the monitored area; • a communication interface (K) by means of which the action (A) is transmitted to a user or to a control unit. [2] Device (100) according to claim 1, wherein • an encoder of the first machine learning model (12) encodes the image data (41) as vectors in a vector space and thereby obtains image embeddings, a decoder of the first machine learning model (12) decodes the vectors into a first text, the first hardware unit (10) outputs the first text as the first character sequence (11) and the hardware unit (30) obtains the first text as the first character sequence (11) via the data interface (D); • an encoder of the second machine learning model (22) encodes the audio data (51) as vectors in a vector space and thereby obtains audio embeddings, a decoder of the second machine learning model (22) decodes the vectors into a second text, the second hardware unit (20) outputs the second text as the second character sequence (11), and the hardware unit (30) obtains the second text as the second character sequence (21) via the data interface (D). [3] Device (100) according to claim 1, wherein • an encoder of the first machine learning model (12) encodes the image data (41) as vectors in a vector space and thereby obtains image embeddings, the first hardware unit (10) outputs the image embeddings as the first character sequence (11) and the hardware unit (30) obtains the image embeddings as the first character sequence (11) via the data interface (D); • an encoder of the second machine learning model (22) encodes the audio data (51) as vectors in a vector space and thereby obtains audio embeddings, the second hardware unit (20) outputs the audio embeddings as the second character sequence (11), and the hardware unit (30) obtains the audio embeddings as the second character sequence (21) via the data interface (D). [4] Device (100) according to one of the preceding claims, wherein actions (A) are stored in the basic model (31) and the basic model (31) selects the action (A) matching this link from the stored actions (A) based on the link. [5] Device (100) according to claim 4, wherein in the case that no suitable action (A) can be selected from the stored actions (A), the basic model (31) determines a new action (A) for this case or prompts the user via a text output to determine a new action (A) and stores the new action (A). [6] Device (100) according to one of the preceding claims, wherein the hardware unit (30) is integrated into a communication network of a vehicle and the communication interface (K) transmits the action (A) to a vehicle control unit for controlling and / or regulating a longitudinal and / or transverse movement of the vehicle and the vehicle control unit executes the action (A). [7] System (200) for monitoring a specific spatial and / or functional area by means of a data interface (D), the system (200) comprising: • at least one image sensor (40) and a first hardware unit (10), wherein the first hardware unit (10) executes a first machine learning model (12) that has been trained on image data to recognize, classify, localize and / or track objects and to convert them into a correspondingly descriptive character sequence, wherein the first hardware unit (10) reads in the image data (41) of the image sensor (40) and outputs a first character sequence (11); • at least one acoustic sensor (50) and a second hardware unit (20), wherein the second hardware unit (20) executes a second machine learning model (22) trained on audio data to detect, classify, localize and / or track acoustic events and to convert them into a correspondingly descriptive character sequence, wherein the second hardware unit (20) reads in the audio data (51) of the acoustic sensor (50) and outputs a second character sequence (21); • a third hardware unit (30) executing a pre-trained basic model (31) which reads in the first character sequence (11) and the second character sequence (21) via the data interface (D), logically links the first character sequence (11) and the second character sequence (21) based on the pre-training, and determines an action (A) based on the link to ensure the safety of persons and objects in the monitored area; • a communication interface (K) by means of which the action (A) is transmitted to a user or a control device. [8] Use of the system (200) according to claim 6 as a driving system of a vehicle, building management system, industrial control system or environmental monitoring system or as a part of one of the above-mentioned systems. [9] Computer-implemented method for monitoring a specific spatial and / or functional area by means of a data interface (D), the method comprising the steps: • with an image sensor (40) obtaining an image from the area (V1); • with an acoustic sensor (50) obtaining an audio recording from the area (V2); • Inputting the image recording into a first machine learning model (12) which has been trained on image data to detect, classify, localize and / or track objects and converting it into a correspondingly descriptive character sequence, and outputting a first character sequence (11) (V3) correspondingly descriptive of the area; • Inputting the audio recording into a second machine learning model (22) trained on audio data to detect, classify, localize and / or track acoustic events and converting them into a correspondingly descriptive character sequence, and outputting a second character sequence (21) correspondingly descriptive of the area (V4); • Inputting the first text (11) and the second text (21) into a hardware unit (30) by means of the data interface (D), which executes a pre-trained basic model (31) which, based on the pre-training, logically links the first character sequence (11) and the second character sequence (21) and, based on the link, determines an action (A) to ensure the safety of persons and objects in the monitored area (V5); • with a communication interface (K) transmitting the action (A) to a user or to a control unit (V6). [10] Computer program for monitoring a specific spatial and / or functional area by means of a data interface (D), the computer program comprising program instructions which cause a hardware unit (30) to carry out the steps of the method according to claim 9 when the hardware unit (30) executes or loads the computer program.
Citation Information
Patent Citations
Method for training a decision algorithm used in a motor vehicle and motor vehicle
DE102015007493A1
Methods for increasing the detection accuracy of a monitoring system
DE102021206618A1
IN-VEHICLE SYSTEM FOR ESTIMATE A SCENE IN A VEHICLE INTERIOR
DE112019000961T5
AUTOMATIC MONITORING SYSTEM FOR INDEPENDENT PERSONS WHO OCCASIONALLY NEED HELP
DE60208166T2
Detection of aggressive behaviour in public transportation
EP3493171A1