Voice-triggered intelligent safety device and system
A voice-triggered safety system addresses the challenge of stopping machinery during emergencies by using sound analysis and neural networks to identify and control machinery, enhancing workplace safety and reducing accidents and costs.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2026-04-02
AI Technical Summary
Industrial accidents involving contact with moving machinery parts are difficult to stop using emergency stop buttons, especially when they are out of reach or in noisy environments, leading to significant costs and fatalities.
A voice-triggered safety system that uses sound collection devices, neural networks for keyword and sentiment analysis, and machine control to identify emergencies, locate affected machinery, and shut it down.
Enables rapid emergency response by identifying worker profiles, machinery location, and controlling machines based on voice inputs, reducing workplace accidents and associated costs.
Smart Images

Figure 2026057467000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to safety systems in industrial environments, and more specifically, to intelligent voice-triggered safety systems.
Background Art
[0002] Industrial accidents in the workplace are a common problem. According to the Bureau of Labor Statistics of the US Department of Labor, in the past few decades, about 5,000 fatal workplace accidents have been recorded in the United States every year. A significant number (more than 500) of these are caused by contact between people and moving parts of machinery. During this time, it can even be difficult to stop the machine using an emergency stop button or call for help. This can occur due to situations where the injured person cannot reach the safety emergency stop button, is unconscious, or other unprecedented situations. According to the National Safety Council, the total cost of workplace accidents in 2021 was $167 billion, and the cost per fatality was $1.3 million.
Summary of the Invention
[0003] In cases where an emergency involving a person and moving machine parts occurs, the physical emergency button is likely to be out of reach. The implementation examples described in this specification include a safety system that can be triggered by voice, identify the related machinery to prevent further injury, and shut it down. Specifically, there are three problems of related technologies to be addressed by the embodiments described in this specification.
[0004] It is necessary to stop the machine without physical or manual operation. Also, in a potentially noisy environment, it is necessary to identify and confirm an actual emergency based on voice or sound input from a microphone. It is also necessary to locate and identify the source of the emergency. For example, when there are multiple facilities in a factory or warehouse, it is necessary to determine which facility should be stopped and shut down.
[0005] Embodiments described herein may include a system capable of detecting emergencies in various manufacturing environments, the system comprising at least one sound collection device, at least one memory for storing collected sound data, at least one device for converting the sound data from analog signals to digital signals, at least one memory containing executable actions by a processor for processing the collected sound data, extracting human voice data from the collected data, extracting ambient sounds from the collected data, generating timestamped sound data based on data collection time, and detecting the presence of specific words in the human voice data. One such implementation may include training a neural network using labeled human voice data, analyzing emotions from the collected voice to determine whether emotions associated with an emergency, such as fear, panic, or anxiety, determining whether an emergency exists using analysis such as keyword detection and sentiment analysis, and transmitting signals to control the affected machine(s) according to a predefined emergency mitigation plan, such as stopping or slowing down the machine(s), or setting alarms to off.
[0006] The embodiment may further include instructions for identifying workers, which include generating a profile using collected human voice data, the profile may function as a “voiceprint”, identifying a group of workers currently working in a specific area based on a work schedule, comparing the generated profile with a database containing the profiles of the group of workers, and calculating the confidence level of worker profile identification.
[0007] Examples may further include instructions to identify the source of the sound, the affected machinery, the location of the affected machinery, and any workers at risk. One such implementation compares the intensity of sound collected by multiple machines.
[0008] Aspects of the present disclosure may include a system for a manufacturing environment, the system comprising at least one sound collector, a memory configured to store sound data collected from at least one sound collector, an analog-to-digital converter configured to convert the stored sound data from analog sound data to digital sound data, and a processor, the processor extracting human voice data from the digital sound data, extracting ambient sound data from the digital sound data, performing word detection on the human voice data, performing sentiment analysis on the extracted human voice data, and, in response to word detection and sentiment analysis indicating an emergency, controlling one or more relevant machines in the manufacturing environment in response to an emergency, the location of one or more relevant machines being derived from the ambient sound data.
[0009] Aspects of the present disclosure may include a method for a manufacturing environment, the method comprising: storing sound data collected from at least one sound collection device; converting the stored sound data from analog sound data to digital sound data; extracting human voice data from the digital sound data; extracting ambient sound data from the digital sound data; performing word detection on the human voice data; performing sentiment analysis on the extracted human voice data; and controlling one or more relevant machines in the manufacturing environment in response to an emergency, based on the analysis of word detection and sentiment analysis indicating an emergency, wherein the location of one or more relevant machines is derived from the ambient sound data.
[0010] Aspects of the present disclosure include a computer program for storing instructions for a manufacturing environment, which may include storing sound data collected from at least one sound collector; converting the stored sound data from analog sound data to digital sound data; extracting human voice data from the digital sound data; extracting ambient sound data from the digital sound data; performing word detection on the human voice data; performing sentiment analysis on the extracted human voice data; and controlling one or more relevant machines in the manufacturing environment in response to an emergency, based on the analysis of word detection and sentiment analysis indicating an emergency, wherein the location of one or more relevant machines is derived from the ambient sound data. The computer program and instructions may be stored in a non-transient computer-readable medium and may be executed by one or more processors.
[0011] Aspects of the present disclosure may include a system for a manufacturing environment, which includes means for storing sound data collected from at least one sound collecting device; means for converting the stored collected sound data from analog sound data to digital sound data; means for extracting human voice data from the digital sound data; means for extracting ambient sound data from the digital sound data; means for performing word detection on the human voice data; means for performing sentiment analysis on the extracted human voice data; and means for controlling one or more relevant machines in the manufacturing environment in response to an emergency, in response to word detection and sentiment analysis indicating an emergency, wherein the location of one or more relevant machines may be derived from the ambient sound data.
[0012] Next, a general architecture for implementing various features of this disclosure will be described with reference to the drawings. The drawings and related descriptions are provided to illustrate embodiments of this disclosure and do not limit the scope of this disclosure. Throughout the drawings, reference numbers are reused to indicate correspondences between referenced elements. [Brief explanation of the drawing]
[0013] [Figure 1] Figure 1 shows the overall workflow of an audio-based emergency detection system in an example. [Figure 2] Figure 2 shows a sound collection component according to an embodiment. [Figure 3] Figure 3 shows the audio signal preprocessing components according to the embodiment. [Figure 4] Figure 4 shows a keyword detection component according to an embodiment. [Figure 5] Figure 5 shows a component for sentiment analysis according to an example. [Figure 6] Figure 6 shows the worker identification component according to the embodiment. [Figure 7] Figure 7 shows the components for sound source identification according to the embodiment. [Figure 8] Figure 8 shows an emergency verification component according to an embodiment. [Figure 9] Figure 9 shows the components corresponding to taking emergency action according to the embodiment. [Figure 10] Figure 10 shows a typical application where all machinery and equipment in the shop are equipped according to the embodiments described herein. [Figure 11] Figure 11 shows an example where the embodiment is not implemented in all machines and equipment. [Figure 12] Figure 12 shows an example of an emergency situation occurring when multiple devices simultaneously detect emergency signals of similar strength. [Figure 13] Figure 13 shows a plurality of machines configured to operate according to the embodiments described herein. [Figure 14] Figure 14 shows an exemplary computing environment with exemplary computer devices suitable for use in several embodiments. [Modes for carrying out the invention]
[0014] The following detailed description provides details of the figures and embodiments of this application. Reference figures between figures and descriptions of redundant elements have been omitted for clarity. Terms used throughout this specification are provided as examples and are not intended to limit them. For example, the use of the term “automatic” may include fully automatic or semi-automatic embodiments with user or administrator control over specific aspects of the embodiment, depending on a desired embodiment for those skilled in the art practicing embodiments of the present invention. Selection may be performed by a user through a user interface or other input means, or through a desired algorithm. The exemplary embodiments described herein may be used individually or in combination, and the functions of the embodiments may be performed by any means according to a desired embodiment.
[0015] Figure 1 shows the overall workflow of sound-based emergency detection according to an embodiment. The embodiment described herein may include eight components, as shown in Figure 1. The first component is a sound collection component 101, which is described in more detail in Figure 2. The second component is a sound signal preprocessing component 102, which is described in more detail in Figure 3. The third component is a keyword detection component 103, which is described in more detail in Figure 4. The fourth component is a sentiment analysis component 104, which is described in more detail in Figure 5. The fifth component is a worker identification component 105, which is described in more detail in Figure 6. The sixth component is a source identification component 106, which is described in more detail in Figure 7. The seventh component is an emergency confirmation configuration component 107, which is described in more detail in Figure 8. The eighth component is an emergency action component 108, which is described in more detail in Figure 9.
[0016] FIG. 2 is a diagram showing the sound collection component 101 according to an embodiment. The sound collection component collects the sound 201 in the environment using a microphone for a preset period of time. The microphone 202 for sound collection can be attached to a machine, adjacent to a machine, or integrated into a wearable device according to a desired implementation. The recorded raw sound data 203 is saved for further analysis.
[0017] FIG. 3 shows the components of the sound signal preprocessing 102 according to an embodiment. This component includes three offline-trained neural networks for detecting background sound 301, human voice 302, and environmental sound 303, respectively. The trained background sound detector 301 detects the working noise 310 in a factory / warehouse. The environmental sound extractor 303 detects environmental sounds 312 such as machine sounds. The voice extractor 302 detects and extracts human voice data 311 from the raw sound data 203.
[0018] In the embodiment, after receiving the raw sound data 203 from the flow of FIG. 2, the sound data is first converted into a digital signal using an analog-to-digital converter (ADC) 304. Next, the background sound is removed / filtered by the trained background sound detector 301 so that the remaining digital signal can be processed by the voice extractor 302, the environmental sound extractor 303, and the data logger 305. The converted and filtered digital signal is separated into human voice 320, environmental sound 321, and timestamped sound data 322.
[0019] Figure 4 shows the keyword detection component 103 according to an embodiment. This component has a neural network that is trained offline using human voice data 311. After receiving the human voice data 320 extracted from the voice extractor 302, the trained neural network converts the voice into text 403. The text 403 is checked against a set of predefined keywords 404, such as "Help! Help!", "Stop!", etc., depending on the desired implementation. The output of this component is whether one or more keywords are present in the voice data (the presence or absence of one or more keywords). The presence or absence of keywords 406 is generated based on a determination from the determination algorithm 405.
[0020] Figure 5 shows the emotion analysis component 104 according to an embodiment. The purpose of this component is to detect emotions related to emergency situations such as fear, anxiety, panic, etc. This component includes a neural network that is trained offline using labeled human voice data from public datasets 501 to form a trained voice recognition neural network. Each voice data is labeled with the emotion(s) associated with the voice snippet at 512, and this is used to train a deep neural network to form an emotion classifier 510. After receiving the human voice data 320 from the human voice extractor 302, the trained emotion classifier 510 outputs the type of emotion 511 expressed in the voice.
[0021] Figure 6 shows a worker identification component 105 according to an embodiment. Two pieces of information, a profiling model 601 and a worker voice profile database 602, are prepared offline. The profiling model 601 is a neural network trained using human voice data 311. The worker voice profile database 602 contains the voice profiles of all workers that can be used as unique identifiers. Such a database can be constructed by a neural network that takes in pre-recorded human voice data 312 of the workers to be identified and outputs the corresponding voice profile / voiceprint 313 for each worker.
[0022] After preprocessing, the extracted human voice data 320 is provided to the voice profiling model 601. The timestamped voice data 322 is used to identify candidate workers and their unique voice profiles by referencing a database of metadata associated with each worker 613, along with metadata for the work 610, such as factory work schedules, badge scan information, and login information. The determination algorithm 612 uses the generated voice profiles 611, candidate workers and their voice profiles from the database to calculate a list of workers with a corresponding confidence level 614, and a list of workers with a confidence level, profile, and metadata 615. Furthermore, if other sensor modules 616 are available to obtain worker locations 617, each identified worker is also associated with its physical location information, as shown in 618.
[0023] Figure 7 shows a sound source identification component 106 according to an embodiment. This component is for identifying the location of machinery and workers in an emergency. Ambient sounds 321 received by the machinery may have different intensities indicating differences in physical distance between the machinery and the worker in the emergency. Each machine that receives sound is assigned a radius based on its acoustic intensity level (level of sound strength). The intersection of the circles from the affected machinery represents the location of the worker at risk. The sound source identification 700 and the corresponding location can be provided accordingly.
[0024] Figure 8 shows an emergency confirmation component 107 according to an embodiment. The emergency (EMG) determination model 800 utilizes information from previous components, including the presence of one or more predefined keywords 406, the presence of an emotion associated with the emergency 511, worker identification with a calculated confidence level 614 (or 615, 616), and sound source location 700. Users can also customize the determination module to emphasize different elements such as sentiment analysis and sound source location according to their desired implementation. For example, it is possible to increase the weight of sentiment analysis in EMG determination so that high levels of emotion (despair / frustration) can be directly determined as EMG positive.
[0025] In 801, the confidence level of the judgment is determined. If there is confidence in the EMG judgment (Yes), the output 802 of this component is whether the recorded sound indicates an emergency. If an emergency is confirmed, information on the affected machine(s) and worker(s) becomes available to take appropriate action. On the other hand, if the confidence level of the EMG judgment is low (No), the EMG verification model 803 is executed to directly verify whether an actual emergency scenario exists (i.e., by asking "Is this a real emergency?"), and at the same time, the response is routed to trigger the sound collection 101 so that the analysis can be run again.
[0026] Figure 9 shows the components corresponding to taking emergency action 108 according to the embodiment. After confirming that an emergency has occurred based on sound data, several actions can be taken as determined by the EMG execution model 900, which may include sending a signal 901, issuing a command to the controller 902 of the affected machine, an electronic notification 903, and raising an alarm 904. The user can also adjust the sensitivity to trigger the level of action to be taken, such as notification only, machine E-stop, power cut-off, etc., according to the desired implementation. Typically, the affected machine can be stopped or slowed down by sending a signal to the machine controller. Electronic notifications can be sent to workers in the factory, and alarms in the affected area can also be set to ON.
[0027] Figure 10 shows a typical application where all machines and equipment in the store are equipped according to the embodiments described herein. If an emergency occurs at machine 6 where worker 6 is present, the embodiments described herein will detect the worker in need of help, stop machine 6, notify surrounding workers, sound a warning alarm, prevent further damage, and provide assistance. This can also be applied to a scenario with worker X and machine Y, provided the machine is equipped with the present invention.
[0028] Figure 11 shows an example where not all machinery and equipment are equipped with the implementation of the embodiment. In particular, all machinery and heavy equipment are equipped according to the embodiment described herein, but the two gates are not equipped. However, the emergency occurred at either worker 6 or worker at gate 1. Even in this case, the embodiment described herein picks up some distant signal, processes the information, and takes action. However, since the gates are not controlled by the embodiment described herein, the emergency action is limited to sounding an alarm, notifying colleagues or supervisors, calling 911, etc., rather than directly stopping the machinery.
[0029] Figure 12 illustrates an example of an emergency occurring when multiple devices simultaneously detect emergency signals of similar strength. In this case, the emergency occurred at worker 6, located at a similar distance from all machines 3, 4, 5, and 6, as shown in the figure. In this scenario, the user can adjust the sensitivity level of the emergency action to stop all machines 3-6 and notify the corresponding workers and supervisors. In extreme cases, all machines can be stopped when an emergency is detected.
[0030] The embodiments described herein enable a rapid response to emergencies when workers are unable to physically stop machinery causing a hazard. This can lead to manufacturers improving workplace safety, saving lives related to worker injury and death, and reducing costs and productivity losses.
[0031] The exemplary embodiments described herein can further utilize work-related information (such as worker schedules and worker profiles) to identify emergencies with high accuracy and facilitate customizable emergency actions to protect workers.
[0032] Furthermore, the embodiments described herein may not only reduce the premiums associated with manufacturers' asset insurance, but also reduce the costs that insurance companies pay for injury and death.
[0033] While the exemplary implementations described herein are geared towards use cases in manufacturing environments, the same systems / solutions can be applied to other industrial sectors and applications involving people, mobile equipment, or rotating machinery, such as conveyor systems, forklifts, robots, AGVs, warehouse cranes, automated truck / ship / airplane docking stations, building escalators and elevators, and construction machinery on site.
[0034] Figure 13 shows a plurality of machines configured to operate according to exemplary embodiments described herein. One or more machines 1321 (e.g., a belt conveyor, an air compressor, a lathe, a forklift, a press, etc.) are configured to perform their corresponding functions and can be communicably coupled to a network 1320 (e.g., a local area network (LAN), a wide area network (WAN)) via corresponding network interfaces of sensor systems installed on the machines 1321, which is connected to a control device 1322 configured to facilitate functions for object recognition. One or more machines 1321 can be associated with sensors or other data acquisition mechanisms, depending on the desired embodiment. The control device 1322 manages a database 1323, which stores historical data collected from sensor systems or data acquisition mechanisms from each machine 1321. In an alternative embodiment, data from the sensor system of machine 1321 can be stored in a dedicated database that takes in data from machine 1321, or in a central repository or central database such as an enterprise resource planning system, and the management device 1322 can access or retrieve the data from the central repository or central database.
[0035] The control device 1322 can also be configured, depending on the desired embodiment, to function as a direct controller of one or more machines 1321 to control the operation of one or more machines 1321, or to send instructions to local controllers of one or more machines 1321 to control one or more machines 1321.
[0036] The sensor system of machine 1321 may include, but is not limited to, any type of sensor to facilitate a desired implementation and provide internal status machine data, such as a gyroscope, accelerometer, vision sensor (e.g., camera, depth camera, infrared sensor, etc.), Global Positioning Satellite System (GPS), thermometer, hygrometer, or any sensor according to the desired implementation. The control device 1322 may also be connected to one or more sound-generating devices (not shown) that monitor the external status of one or more machines 1321 by collecting sound data, as described herein.
[0037] Figure 14 shows an exemplary computing environment having exemplary computer devices suitable for use in several exemplary implementations, such as a control device 1322 for facilitating the functions of each robot. The computer device 1405 of the computing environment 1400 may include one or more processing units, cores, or processors 1410, memory 1415 (e.g., RAM, ROM, and / or similar), internal storage 1420 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or I / O interfaces 1425, any of which may be coupled on a communication mechanism or bus 1430 for communicating information, or embedded in the computer device 1405. The I / O interface 1425 may also be configured, depending on the desired implementation, to receive images from a camera or provide images to a projector or display.
[0038] Computer device 1405 may be communicatively coupled to an input / user interface 1435 and an output device / interface 1440. Either or both of the input / user interface 1435 and the output device / interface 1440 may be wired or wireless interfaces and may be detachable. The input / user interface 1435 may include any physical or virtual devices, components, sensors, or interfaces that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Braille, motion sensors, optical readers, and / or similar). The output device / interface 1440 may include displays, televisions, monitors, printers, speakers, Braille, and the like. In some exemplary implementations, the input / user interface 1435 and the output device / interface 1440 may be embedded in or physically coupled to computer device 1405. In other exemplary implementations, other computer devices may function as or provide the functions of the input / user interface 1435 and output device / interface 1440 of computer device 1405.
[0039] Examples of computer devices 1405 include, but are not limited to, highly mobile devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by people and animals), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, etc., having one or more processors embedded therein and / or coupled thereto).
[0040] Computer device 1405 may be communicatively coupled (for example, via I / O interface 1425) to external storage 1445 and network 1450 for communication with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 1405 or any connected computer device may function, provide services, or be referred to as a server, client, thin server, general machine, special-purpose machine, or other label.
[0041] The I / O interface 1425 may include, but is not limited to, wired and / or wireless interfaces using any communication or I / O protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMAX, modem, cellular network protocol, etc.) for communicating information with at least all connected components, devices, and networks within the computing environment 1400. The network 1450 may be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).
[0042] Computer device 1405 may use and / or communicate using computer-usable media or computer-readable media, including transient media and non-transient media. Transient media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transient media include magnetic media (disks, tapes, etc.), optical media (CD-ROMs, digital video discs, Blu-ray discs, etc.), solid-state media (RAM, ROMs, flash memory, solid-state storage, etc.), and other non-volatile storage or memory.
[0043] The computer device 1405 can be used to implement a technology, method, application, process, or computer executable instruction in several exemplary computing environments. The computer executable instruction may be obtained from a transient medium, stored in a non-transient medium, and retrieved from a non-transient medium. The executable instruction may originate from one or more programming languages, scripting languages, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).
[0044] The processor(s) 1410 can run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications can be deployed, including logical units 1460, application programming interface (API) units 1465, input units 1470, output units 1475, and inter-unit communication mechanisms 1495 for different units to communicate with each other, with the OS, and with other applications (not shown). The units and elements described may vary in design, function, configuration, or implementation, and are not limited to the description provided. The processor(s) 1410 may take the form of a hardware processor, such as a central processing unit (CPU), or a combination of hardware and software units.
[0045] In some exemplary implementations, once information or execution instructions are received by the API unit 1465, they may be transmitted to one or more other units (e.g., a logical unit 1460, an input unit 1470, and an output unit 1475). In some embodiments, the logical unit 1460 may be configured to control the flow of information between units and to direct the services provided by the API unit 1465, the input unit 1470, and the output unit 1475, in some exemplary implementations described above. For example, one or more processes or implementation flows may be controlled by the logical unit 1460 alone or in conjunction with the API unit 1465. The input unit 1470 may be configured to take input for the computation described in the exemplary implementation, and the output unit 1475 may be configured to provide output based on the computation described in the exemplary implementation.
[0046] The memory 1415 can be configured to store sound data from at least one sound collection device, as disclosed in the environment of Figure 13. The stored sound data can be converted from analog sound data to digital sound data by an analog-to-digital converter (not shown).
[0047] The processor(s) 1410 may be configured to perform a method or computer instruction that includes: extracting human voice data 320 from digital sound data; extracting ambient sound data 321 from digital sound data; performing word detection 103 on the human voice data 320; performing sentiment analysis 104 on the extracted human voice data 320; and controlling one or more related machines in the manufacturing environment in response to an emergency (Figure 9), based on the analysis of word detection and sentiment analysis indicating an emergency (800 to 802 in Figure 8). 8) Controlling one or more related machines in the manufacturing environment in response to an emergency (Figure 9), where the location of one or more related machines is derived from ambient sound data (Figure 7).
[0048] The processor(s) 1410 may be configured to perform the methods or instructions described above, and further include identifying a worker from the human voice data 320 based on a work schedule (e.g., metadata 610), a timestamp 322 applied to the human voice data 320, and one or more worker profiles (e.g., from database 602) constructed from the execution of a neural network on previously collected human voice data.
[0049] The processor(s) 1410 may be configured to perform the methods or instructions described above, and further include identifying one of the workers currently at risk based on a position radius derived from the acoustic intensity of one or more associated machines, as described with respect to Figure 7.
[0050] The processor(s) 1410 may be configured to perform the methods or instructions described above, and further include sending a notification to an identified worker currently at risk, as described with respect to Figures 7 and 9.
[0051] The processor(s) 1410 can be configured to perform the methods or instructions described above, and further include applying timestamps to ambient sound data and human voice data based on the data acquisition time from at least one sound collection device, as shown in Figures 305 and 322.
[0052] The processor(s) 1410 may be configured to perform the method or instruction described above, where the performance of sentiment analysis is carried out by a sentiment classifier constructed from a neural network trained on a dataset of sounds and corresponding emotions, as shown in Figure 5.
[0053] Some parts of the detailed description are presented in terms of symbolic representations of algorithms and computer operations. These algorithmic descriptions and symbolic representations are means used by those skilled in the field of data processing technology to convey the essence of the innovation. An algorithm is a set of defined steps that lead to a desired final state or result. In the examples, the steps performed require a visible amount of physical operation to achieve the visible result.
[0054] Unless otherwise stated, as will be evident from the discussions, discussions throughout this specification using terms such as “processing,” “calculation,” “calculation,” “determination,” and “display” may include the operations and processes of a computer system or other information processing device that manipulate and convert data represented as physical (electronic) quantities in the registers and memory of a computer system into other data similarly represented as physical quantities in the memory or registers of a computer system or other information storage, transmission, or display device.
[0055] Exemplary embodiments also relate to apparatus for performing the operations described herein. This apparatus may be specifically configured for a particular purpose and may include one or more general-purpose computers that are selectively invoked or reconfigured by one or more computer programs. Such computer programs may be stored on computer-readable media such as computer-readable storage media and computer-readable signal media. Computer-readable storage media include, but are not limited to, tangible media such as optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices, and drives, as well as other types of tangible or non-transient media suitable for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and representations presented herein are not inherently related to any particular computer or other apparatus. Computer programs may include purely software implementations that include instructions for performing the operations of a desired implementation.
[0056] Various general-purpose systems may be used with the programs and modules according to the embodiments described herein, or it may be convenient to construct more specialized devices for performing the desired method steps. Furthermore, the embodiments are not described with reference to any particular programming language. It will be understood that various programming languages can be used to implement the techniques of the embodiments described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), a processor, or a controller.
[0057] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various embodiments of the exemplary implementations may be implemented using circuit and logic devices (hardware), while other embodiments, when performed by a processor, may be implemented using instructions stored on a machine-readable medium (software) that cause the processor to perform the method for performing the implementation of this application. Furthermore, some exemplary implementations of this application may be performed by hardware alone, while other exemplary implementations may be performed by software alone. Moreover, the various functions described may be performed by a single unit or may span a number of components in any number of ways. When performed by software, the method may be performed by a processor such as a general-purpose computer based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in a compressed and / or encrypted format.
[0058] Furthermore, other embodiments of the present application will be apparent to those skilled in the art from the consideration of the present specification and the practice of the present technology. Various aspects and / or components of the exemplary embodiments described herein can be used individually or in any combination. This specification and the exemplary embodiments are intended to be for illustrative purposes only, and the true scope and spirit of the present application are shown by the following claims.
Claims
1. A system for the manufacturing environment. At least one sound collection device, A memory configured to store sound data collected from at least one of the sound collection devices, An analog-to-digital converter configured to convert the stored sound data from analog sound data to digital sound data, Processor and It has, The processor is configured to extract human voice data from the digital sound data, extract ambient sound data from the digital sound data, perform word detection on the human voice data, perform sentiment analysis on the extracted human voice data, and, in response to word detection and sentiment analysis indicating an emergency, control one or more relevant machines in the manufacturing environment in response to the emergency, wherein the location of one or more of the relevant machines is derived from the ambient sound data.
2. In the system described in claim 1, The processor is configured to identify a worker from a person's voice data based on a work schedule, a timestamp applied to the person's voice data, and one or more worker profiles constructed from the execution of a neural network on previously collected voice data of the person.
3. In the system described in claim 2, The processor is further configured to identify the worker currently at risk based on a positional radius derived from the acoustic intensity of one or more associated machines.
4. In the system for the manufacturing environment described in claim 3, The processor is configured to send notifications to identified workers who are currently at risk.
5. In the system described in claim 1, The processor is configured to apply a timestamp to the ambient sound data and the human voice data based on the data acquisition time from at least one of the sound collection devices.
6. In the system described in claim 1, The aforementioned sentiment analysis is performed by a system in which a sentiment classifier is constructed from a neural network trained on a dataset of speech and corresponding emotions.
7. A method for the manufacturing environment, To store sound data collected from at least one sound collection device, Converting the stored audio data from analog audio data to digital audio data, Extracting human voice data from the aforementioned digital audio data, Extracting ambient sound data from the aforementioned digital sound data, Performing word detection on the aforementioned person's voice data, Performing sentiment analysis on the extracted voice data of the person, A method comprising detecting a word indicating an emergency and performing sentiment analysis, and controlling one or more associated machines in a manufacturing environment in response to the emergency, wherein the location of one or more of the associated machines is derived from ambient sound data.
8. In the method according to claim 7, A method for identifying a worker from a person's voice data, based on a work schedule, a timestamp applied to the person's voice data, and one or more worker profiles constructed from the execution of a neural network on previously collected voice data of the person.
9. In the method according to claim 8, A method further comprising identifying the worker currently at risk based on a positional radius obtained from the acoustic intensity of one or more associated machines.
10. In the method according to claim 9, A method further comprising sending a notification to any of the workers currently at risk who have been identified.
11. In the method according to claim 7, A method further comprising applying a timestamp to the ambient sound data and the human voice data based on the data collection time from at least one of the sound collection devices.
12. In the method according to claim 7, The aforementioned sentiment analysis is performed by a sentiment classifier constructed from a neural network trained on a dataset of speech and corresponding emotions.
13. In a non-transient, computer-readable medium storing instructions for a manufacturing environment, The aforementioned instructions include storing sound data collected from at least one sound collection device, Converting the stored audio data from analog audio data to digital audio data, Extracting human voice data from the aforementioned digital audio data, Extracting ambient sound data from the aforementioned digital sound data, Performing word detection on the aforementioned person's voice data, Performing sentiment analysis on the extracted voice data of the person, A non-transient computer-readable medium in which the location of one or more of the associated machines is derived from ambient sound data, comprising: word detection and sentiment analysis indicating an emergency; and control of one or more associated machines in a manufacturing environment in response to the emergency.
14. In the non-transient computer-readable medium described in claim 13, The instructions are a non-transient, computer-readable medium that includes identifying a worker from the person's voice data based on a work schedule, a timestamp applied to the person's voice data, and one or more worker profiles constructed from the execution of a neural network on the person's voice data previously collected.
15. In the non-transient computer-readable medium described in claim 14, The instructions are a non-transient, computer-readable medium that further includes identifying the worker currently at risk based on a positional radius derived from the acoustic intensity of one or more associated machines.
16. In the non-transient computer-readable medium described in claim 15, The instructions are provided in a non-transient, computer-readable medium, which further includes sending a notification to any of the identified workers currently at risk.
17. In the non-transient computer-readable medium described in claim 13, The instructions further include applying a timestamp to the ambient sound data and the human voice data based on the data collection time from at least one of the sound collection devices in a non-transient, computer-readable medium.
18. In the non-transient computer-readable medium described in claim 13, The aforementioned sentiment analysis is performed in a non-transient, computer-readable medium by a sentiment classifier constructed from a neural network trained on a dataset of speech and corresponding emotions.