Information processing device, information processing method, and program

The information processing device improves task classification accuracy by associating sound data with spectrogram images and using machine learning to prioritize similar features, enhancing the reliability of task identification.

JP7790990B2Active Publication Date: 2025-12-23NS SOLUTIONS CORPORATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2022009679
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-01-25
Publication Date
2025-12-23
Estimated Expiration
2042-01-25

AI Technical Summary

Technical Problem

Conventional technologies struggle to achieve accurate classification of tasks performed by workers and equipment in a work environment.

Method used

An information processing device that associates sound data with spectrogram images and uses machine learning to classify tasks by assigning higher weights to data with features similar to pre-registered information, incorporating certainty information from a recognizer.

Benefits of technology

Enhances the accuracy of task classification by prioritizing data with higher similarity to pre-registered features, improving the reliability of task identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007790990000002
    Figure 0007790990000002
  • Figure 0007790990000003
    Figure 0007790990000003
  • Figure 0007790990000004
    Figure 0007790990000004
Patent Text Reader

Abstract

To more accurately classify work, which is conducted by a subject to be a target.SOLUTION: An information processor includes: association means for associating data on sound based on a sound collection result by a sound collecting device worn by a predetermined subject with information on a detection object indicated by the sound as incidental information on the basis of an analysis result of the sound; weighting means for setting weight for a series of the data in which the incidental information is associated so that priority is given to data in which a feature based on information registered in advance for work and the incidental information on a feature having higher similarity are associated, for each work which is a candidate of a classification object; and classification means for classifying the work conducted by the subject on the basis of the data in which the weight is set for each work which is the candidate of the classification object.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] Conventionally, there have been known techniques for determining the work being performed by a worker or equipment (e.g., a work vehicle) by using information corresponding to the observation results of the status of the work environment, such as moving images or still images (hereinafter, these may be collectively referred to as images) corresponding to the imaging results of the work environment. In recent years, various techniques have been considered as an example of such a technique, in which a trained model constructed in advance based on machine learning is used to determine the work being performed by a worker or equipment. For example, Patent Document 1 discloses an example of a technique for determining the work being performed by a work vehicle by using image data (hereinafter, also referred to as image data) corresponding to the imaging results of a camera attached to the work vehicle. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2020-4096 Summary of the Invention [Problem to be solved by the invention]

[0004] On the other hand, conventional technologies may not always be able to achieve the required accuracy when classifying tasks performed by task subjects such as workers and equipment (e.g., work vehicles) in a work environment. In light of this, there is a demand for technology that enables more accurate classification of tasks performed by target subjects (e.g., workers and equipment).

[0005] In view of the above problems, the present invention aims to enable more accurate classification of the work being performed by a target entity. [Means for solving the problem]

[0006] The information processing device according to the present invention comprises an association means for associating, as incidental information, information on a detection target indicated by sound with respect to sound data based on a sound collection result by a sound collection device worn by a predetermined subject, based on an analysis result of the sound; a weighting means for setting weights on a series of data associated with the incidental information so that, for each task that is a candidate for classification, data associated with the incidental information relating to a feature more similar to a feature based on information pre-registered for the task is given higher priority; and a classification means for classifying tasks being performed by the subject based on the data with the weight set for each task that is a candidate for classification, wherein the weighting means associates, as incidental information, information on a detection target indicated by the sound with the data corresponding to the sound, based on a result of analysis of a spectrogram image converted from sound according to the sound collection result by the sound collection device, and for each task that is a candidate for classification, data associated with the incidental information relating to a feature more similar to a feature based on one or more detection targets pre-registered for the task is given higher priority. Confidence Information The association means assigns weights to the series of data associated with the additional information so that the data associated with the additional information is given higher priority, and the association means inputs the spectrogram image converted from the sound based on the sound collection results by the sound collection device to a recognizer constructed based on machine learning, and associates certainty information output from the recognizer, which indicates the likelihood that the object indicated by the sound is the detection target, as the additional information, with the data corresponding to the sound. [Effects of the Invention]

[0007] According to the present invention, it is possible to more accurately classify the work being performed by a target entity. [Brief explanation of the drawings]

[0008] [Figure 1]FIG. 1 illustrates an example of a system configuration of an information processing system. [Figure 2] FIG. 1 illustrates an example of a hardware configuration of an information processing device. [Figure 3] FIG. 2 is a functional block diagram illustrating an example of a functional configuration of the information processing system. [Figure 4] FIG. 10 is a diagram illustrating an example of a process related to determination based on a feature amount. [Figure 5] FIG. 10 is a diagram illustrating an example of a process for setting weights to data. [Figure 6] 10 is a flowchart illustrating an example of processing of the information processing system. [Figure 7] FIG. 10 is a diagram showing an example of analysis processing for image data of a moving image. [Figure 8] FIG. 10 is a diagram for explaining another example of processing relating to determination based on feature amounts. [Figure 9] FIG. 10 is a diagram showing an example of a spectrogram image. [Figure 10] 10 is a flowchart showing another example of processing of the information processing system. DETAILED DESCRIPTION OF THE INVENTION

[0009] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0010] <System configuration> An example of a system configuration of an information processing system according to an embodiment of the present disclosure will be described with reference to FIG. 1. The information processing system 1 according to this embodiment includes a server device 100, one or more terminal devices 200, and a wearable device 300. Note that the terminal devices 200a and 200b shown in FIG. 1 each represent an example of the terminal device 200. In the following description, when there is no need to particularly distinguish between the terminal devices 200a and 200b, they will simply be referred to as the terminal device 200. Furthermore, the wearable device 300 is a terminal device worn by a user and is individually provided for each user to be observed.

[0011] In addition, an observation device 310 that observes the situation around the user wearing the wearable device 300 is supported on the wearable device 300. With this configuration, the observation device 310 is used in a state where it is worn by the user via the wearable device 300. The observation device 310 may be realized, for example, by an imaging device that captures images of the surrounding environment and outputs image data (e.g., still images or moving images) corresponding to the imaging results to a predetermined output destination. As another example, the observation device 310 may be realized by a sound collection device that collects sounds (e.g., environmental sounds, voices, etc.) propagating through the surrounding space and outputs sound data corresponding to the collected sounds to a predetermined output destination. Thus, the type of observation device 310 is not particularly limited, as long as it can observe the surrounding environment in a manner corresponding to the type of sense (modality) that humans can perceive.

[0012] The type of device applied as wearable device 300 is not particularly limited as long as it is a device that can be worn by a user and is configured to support observation device 310. As a specific example, a so-called glasses-type device may be applied as wearable device 300. As another example, observation device 310 itself may be configured as wearable device 300. In this case, observation device 310 is used in a state where observation device 310 itself is worn by a user via a support member such as a belt.

[0013] In this embodiment, in order to more clearly explain the characteristics of the information processing system 1, it is assumed that the user to be observed is a single user indicated as user U1, and that the wearable device 300 used is a single device worn by the user U1. It is also assumed that the observation device 310 is an imaging device, such as a video camera, that can output data based on the observation results of the surrounding situation, including images corresponding to the imaging results and sound from the sound collection results.

[0014] The server device 100, each terminal device 200, and the wearable device 300 are connected via a network N1 so as to be able to transmit and receive information to and from each other. The type of network N1 is not particularly limited. As a specific example, network N1 may be configured by the Internet, a dedicated line, a local area network (LAN), a wide area network (WAN), or the like. Network N1 may be configured by a wired network or a wireless network such as a network based on communication standards such as 5G, LTE (Long Term Evolution), and Wi-Fi (registered trademark). Network N1 may include multiple networks, and some of the networks may be of a different type from the other networks. It is sufficient that communication between the various information processing devices described above is logically established, and physically, communication between the various information processing devices may be relayed by another communication device or the like.

[0015] The terminal device 200 serves as an interface for receiving inputs (e.g., various instructions) from a user and presenting various information (e.g., feedback, etc.) to the user. As a specific example, the terminal device 200 may receive data from the server device 100 (described later) via a network, and present information based on the data to the user via a predetermined output device (e.g., a display, etc.). Furthermore, the terminal device 200 may recognize instructions from a user based on an operation received from the user via a predetermined input device (e.g., a touch panel, etc.), and transmit information corresponding to the instructions to the server device 100 via the network. This enables the server device 100 to recognize instructions from the user and execute processing according to the instructions. The terminal device 200 can be realized by an information processing device having a communication function, such as a so-called smartphone, a tablet terminal, or a PC (Personal Computer).

[0016] The server device 100 provides various functions to assist a user (hereinafter simply referred to as the administrator) in managing and analyzing the status of work performed by workers to be managed. For example, the server device 100 acquires data from a wearable device 300 worn by a worker to be managed (e.g., user U1) based on the results of observation of the worker's surroundings by an observation device 310 supported by the wearable device 300. The server device 100 analyzes the acquired data and, based on the results of the analysis, classifies the work performed by the worker wearing the wearable device 300 into one of a series of tasks that are candidates for classification. As a specific example, the server device 100 may determine the similarity between feature values ​​extracted from the acquired data and feature values ​​set for each of the series of tasks that are candidates for classification, and then classify the work performed by the worker based on the results of the determination. The above-mentioned functions of the server device 100 will be described in detail later.

[0017] Note that the configuration shown in FIG. 1 is merely an example and does not necessarily limit the system configuration of the information processing system 1 according to this embodiment. As a specific example, the server device 100 may play the role of the terminal device 200. That is, the server device 100 itself may accept input of various information from a user and present various information to the user. Furthermore, a component equivalent to the server device 100 may be realized by multiple devices working together. As a specific example, a component equivalent to the server device 100 may be realized as a so-called cloud service. In this case, the cloud service may be realized by multiple server devices working together.

[0018] An example of the system configuration of the information processing system according to an embodiment of the present disclosure has been described above with reference to FIG.

[0019] <Hardware configuration> 2, an example of a hardware configuration of an information processing device 900 applicable as various devices (e.g., the server device 100, the terminal device 200, and the wearable device 300) constituting the information processing system 1 according to the present embodiment shown in FIG. 1 will be described. The information processing device 900 includes a central processing unit (CPU) 910, a read-only memory (ROM) 920, a random access memory (RAM) 930, an auxiliary storage device 940, and a network I / F 970. The information processing device 900 may also include at least one of an output device 950 and an input device 960. The CPU 910, the ROM 920, the RAM 930, the auxiliary storage device 940, the output device 950, the input device 960, and the network I / F 970 are connected to one another via a bus 980.

[0020] The CPU 910 is a central processing unit that controls various operations of the information processing device 900. For example, the CPU 910 may control the operation of the entire information processing device 900. The ROM 920 stores control programs, boot programs, and the like that can be executed by the CPU 910. The RAM 930 is the main storage memory of the CPU 910, and is used as a work area or a temporary storage area for expanding various programs.

[0021] The auxiliary storage device 940 stores various data and programs. The auxiliary storage device 940 is realized by a storage device capable of temporarily or permanently storing various data, such as a hard disk drive (HDD) or a nonvolatile memory such as a solid state drive (SSD).

[0022] The output device 950 is a device that outputs various types of information and is used to present various types of information to a user. For example, the output device 950 may be realized by a display device such as a display, and may present information to a user by displaying various types of display information. As another example, the output device 950 may be realized by an audio output device that outputs sounds such as voices and electronic sounds, and may present information to a user by outputting sounds such as voices and telegrams. In this way, the device used as the output device 950 may be changed as appropriate depending on the medium used to present information to a user. The output device 950 corresponds to an example of an "output unit" used to present various types of information.

[0023] The input device 960 is used to receive various instructions from the user. For example, the input device 960 may include an input device such as a mouse, a keyboard, or a touch panel. As another example, the input device 960 may include a sound collection device such as a microphone, and collect voices uttered by the user. In this case, various analysis processes such as acoustic analysis and natural language processing may be performed on the collected voices, and the contents of the voices may be recognized as instructions from the user. In this way, the device applied as the input device 960 may be changed as appropriate depending on the method for recognizing instructions from the user. Furthermore, multiple types of devices may be applied as the input device 960.

[0024] The network I / F 970 is used for communication with external devices via a network. The device used as the network I / F 970 may be changed as appropriate depending on the type of communication path and the communication method used.

[0025] The program for the information processing device 900 may be provided to the information processing device 900 by a recording medium such as a CD-ROM, or may be downloaded via a network, etc. When the program for the information processing device 900 is provided by a recording medium, the program recorded on the recording medium is installed in the auxiliary storage device 940 by setting the recording medium in a predetermined drive device.

[0026] 2 is merely an example and does not necessarily limit the hardware configuration of the information processing device that constitutes the information processing system 1 according to this embodiment. As a specific example, some components such as the input device 960 and the output device 950 may not be included. As another example, components according to the functions realized by the information processing device 900 may be added as appropriate.

[0027] An example of the hardware configuration of the information processing device 900 applicable to various devices constituting the information processing system 1 according to this embodiment shown in FIG. 1 has been described above with reference to FIG.

[0028] <Functional configuration> 3, an example of the functional configuration of the information processing system 1 according to this embodiment will be described, focusing particularly on the configuration of the server device 100. The server device 100 includes a communication unit 101, an input / output control unit 102, a data analysis unit 103, a similarity determination unit 106, a weighting processing unit 107, a classification unit 108, and a storage unit 110.

[0029] The communication unit 101 is a communication interface that allows each component of the server device 100 to transmit and receive information to and from other devices (e.g., terminal device 200) via the network N1. The communication unit 101 can be realized, for example, by the network I / F 970. In the following description, when each component of the server device 100 transmits and receives information to and from other devices, it is assumed that the information is transmitted and received via the communication unit 101 unless otherwise specified.

[0030] The storage unit 110 schematically shows a storage area for storing various data, various programs, etc. For example, the storage unit 110 may store data and programs for each component of the server device 100 to execute processing. The storage unit 110 may also store data transmitted from the wearable device 300 (for example, image data or acoustic data corresponding to the observation results of the observation device 310). The storage unit 110 may also store data generated in the process of analyzing the above data by the data analysis unit 103, data generated according to the results of the analysis, etc. The storage unit 110 may also store data used for various determinations by the similarity determination unit 106, which will be described later.

[0031] The input / output control unit 102 performs various processes related to presenting various information to a user (e.g., an administrator) and accepting input of information (e.g., instructions, etc.) from the user. For example, the input / output control unit 102 may perform processes related to presenting a predetermined UI via the terminal device 200 and processes related to accepting input via the UI. This enables the server device 100 to recognize instructions from the user and present the results of processing in accordance with the instructions to the user.

[0032] The data analysis unit 103 acquires data based on the observation results of the surrounding situation of the worker wearing the wearable device 300 by the observation device 310 supported by the wearable device 300, and performs various analyses on the data. Note that the method by which the data analysis unit 103 acquires the data is not particularly limited as long as it can acquire the data. For example, the data analysis unit 103 may receive the data from the wearable device 300. As another example, the data analysis unit 103 may acquire the data by referring to a predetermined storage area in which the data transmitted from the wearable device 300 is stored. The data analysis unit 103 according to this embodiment includes a feature extraction unit 104 and an additional processing unit 105 .

[0033] The feature extraction unit 104 performs a predetermined analysis process on data based on the observation results of the worker's surroundings by the observation device 310, and extracts information indicating the characteristics of the observed surroundings of the worker as features.

[0034] For example, the feature extraction unit 104 may analyze image data (e.g., still images or video images) corresponding to the image capture results of an imaging device applied as the observation device 310. In this case, the feature extraction unit 104 may recognize an object captured as a subject in the image by performing a desired analysis, such as image analysis, on the image corresponding to the image capture results of the imaging device, and extract information corresponding to the recognition result (e.g., text information indicating the object) as a feature. As another example, the feature extraction unit 104 may perform a desired analysis on the image to extract, as a feature, certainty information indicating the likelihood that the object captured as a subject in the image is a predetermined detection target. Furthermore, when analyzing video data, the feature extraction unit 104 may individually extract the feature from a still image corresponding to each frame of a series of frames constituting the video. As another example, the feature extraction unit 104 may analyze sound data corresponding to the sound collection results of a sound collection device applied as the observation device 310. In this case, the feature extraction unit 104 may perform a desired analysis, such as acoustic analysis, on the sound corresponding to the sound collection results of the sound collection device to recognize an object indicated by the sound (e.g., environmental sound, voice, etc.) and extract information corresponding to the recognition result (e.g., text information indicating the object) as a feature. As another example, the feature extraction unit 104 may perform a desired analysis on the sound to extract, as a feature, certainty information indicating the likelihood that the object indicated by the sound is a predetermined detection object. The feature extraction unit 104 may also divide a series of sounds into predetermined periods and individually extract the feature from each of the sounds for each period.

[0035] In addition, the feature extraction unit 104 may apply a trained model (such as a classifier or recognizer) constructed based on so-called machine learning to extract the above-mentioned features indicating the characteristics of the observed surrounding situation of the worker from data based on the observation results of the surrounding situation of the worker by the observation device 310.

[0036] For example, when image data is the subject of analysis processing, a trained model can be applied that is constructed based on machine learning using a pair of an image and information about the subject in the image as training data, so that when an image is input, information indicating the subject in the image is output. Furthermore, when acoustic data is the subject of analysis processing, a trained model constructed based on machine learning using a pair of acoustic data and information related to an object indicated by the acoustic data as training data may be applied, and the trained model is constructed so as to output information related to the object indicated by the acoustic data when the acoustic data is input. As another example, the acoustic data may be converted into a spectrogram, and the image of the spectrogram may be used as the subject of analysis. In this case, a trained model constructed based on machine learning using a pair of an image of a spectrogram and information related to an object indicated by the acoustic data corresponding to the spectrogram, and the trained model is constructed so as to output information related to the object indicated by the acoustic data when the image of the spectrogram converted from the acoustic data is displayed. An example of using a trained model will be described in detail later.

[0037] The additional processing unit 105 associates the feature amounts extracted from the data to be analyzed by the feature amount extraction unit 104 as additional information with the data to be analyzed. As a specific example, the additional processing unit 105 may associate the feature amounts with the data to be analyzed as additional information by so-called tagging processing. Of course, as long as it is possible to associate the feature amounts extracted from the data to be analyzed as additional information with the data to be analyzed, the method is not particularly limited.

[0038] The similarity determination unit 106 determines the similarity between the feature associated with the data and the feature defined for each task that is a candidate for classification, for data that has been subjected to feature extraction and association by the data analysis unit 103. The similarity determined by the similarity determination unit 106 indicates the similarity between the features of the situation around the worker indicated by the target data (i.e., the situation around the observed worker) and the features of the situation around the worker that is expected when the task that is a candidate for classification is being performed.

[0039] An example of the processing performed by the similarity determination unit 106 will now be described in more detail with reference to Fig. 4, focusing on a case where the target data is image data corresponding to the image capture result by an imaging device. Fig. 4 is a diagram outlining a process for determining the similarity between the observed surroundings of a worker and the characteristics of the surroundings of the worker that are expected when a task that is a candidate for classification is being performed, as an example of a process for determination using features extracted from an image. Note that in the example shown in Fig. 4, a trained model is used to extract features from the target image, and certainty information indicating the likelihood that the object captured as a subject in the image is a predetermined detection target is extracted as the feature.

[0040] First, the flow shown on the left side of the example shown in Fig. 4 will be described. In the example shown in Fig. 4, the flow shown on the left side shows a processing flow related to the extraction of feature amounts from images corresponding to the image capture results of the situation around the worker taken by an imaging device. Specifically, by inputting still images corresponding to each of a series of frames constituting a moving image into a trained model, certainty factor information indicating the likelihood that the object captured as a subject in the still images (frames) is the predetermined detection target is extracted.

[0041] For example, in the example shown in FIG. 4, a still image corresponding to a frame capturing a stepladder and a worker's hand is input to the trained model. Therefore, the certainty information is set such that the probability that the object captured as a subject in the still image is a "hand" and a "stepladder" are higher, and the probability that the object is a detection object that is not actually captured as a subject, such as a "dog," is lower. Thus, the certainty information is set to the probability that the observed object is the candidate for each of a series of candidates defined as the detection target. In other words, if there are 1,000 types of detection target candidates, 1,000-dimensional information is output as the certainty information, in which the probability of each of the 1,000 types of candidates is set. It should be noted that various events can be set as detection target candidates, not limited to objects such as "hands" and "stepladders," as long as they are observable by a desired observation device, such as actions such as "cutting" and "chopping" and environmental sounds such as the rustling of grass and trees.

[0042] In the example shown in FIG. 4, a feature vector is extracted based on the probability that the observed object is each of a series of candidates defined as the detection target, based on the confidence information output from the trained model. For example, if there are 1,000 types of detection target candidates, the extracted feature vector will be a 1,000-dimensional vector. A trained model constructed based on machine learning may also be applied to extract the feature vector from the confidence information. A model called "word2vec" is a commonly known example of such a trained model. Hereinafter, the feature vector will also be referred to as a "word vector" for convenience. A word vector based on features (for example, confidence information) extracted from data corresponding to observation results by the observation device 310 (for example, an imaging device) corresponds to an example of a "second feature vector."

[0043] Next, the flow shown on the right side of the example shown in Fig. 4 will be described. In the example shown in Fig. 4, the flow shown on the right side shows a processing flow related to the extraction of feature amounts for each task that is a candidate for classification. In the information processing system according to this embodiment, based on the observation results of the situation around the worker, the task being performed by the worker is classified into one of a series of tasks that have been predefined as candidates for classification. Therefore, for each of the series of tasks that have been predefined as candidates for classification, information about objects that can be observed when the task is being performed (in other words, objects that are highly relevant to the task) is predefined as a task event.

[0044] 4, the names of tools observed when the target work is performed, the names of actions observed when the work is performed, and the names of sound effects (in other words, environmental sounds) observed when the work is performed are defined as work events. Here, focusing on the case of "pruning" work, a specific example of information defined as a work event will be described. As a specific example, in the task of "pruning," tools used in the task (for example, objects reflected in the image as subjects) include "scissors," "saw," "work gloves," "hands," and "branches." Therefore, information indicating these tools is defined as tool names. Furthermore, the task involves actions such as "cutting," "chopping," and "holding." Therefore, information indicating these actions is defined as action names. Furthermore, during the performance of the task, sounds generated by the actions of "cutting" and "chopping," as well as "plant sounds," can be observed. Therefore, information indicating these sounds is defined as sound effect names.

[0045] The method for defining work events is not particularly limited. As a specific example, a manager may define work events corresponding to each work in consideration of the characteristics and implementation status of each work. As another example, work events corresponding to each work may be defined based on observation results of the implementation status of each work. Furthermore, already defined work events may be updated as appropriate depending on the situation at the time. As a specific example, work events corresponding to at least some of the work may be updated based on instructions from a manager. As another example, information according to the implementation status of at least some of the work may be fed back to the work events corresponding to the work, thereby updating the work events.

[0046] Based on the above assumptions, for each task defined as a candidate for classification, a word vector is extracted based on the task event defined for that task. For example, in the case of the "pruning" task shown in Figure 4, a word vector is extracted that includes, as elements, information defined in the corresponding task event, such as the tool name, task name, and sound effect name. The word vector extracted based on the work event corresponds to an example of a "first feature vector."

[0047] Then, the similarity determination unit 106 determines the similarity between a word vector based on the confidence information extracted from an image corresponding to the imaging results of the worker's surroundings and a word vector extracted based on the task event for each task that is a candidate for classification. Note that the method for determining the similarity of these word vectors is not particularly limited as long as it is possible to determine the similarity between multiple vectors. As a specific example, the similarity determination unit 106 may determine the similarity between the two word vectors by calculating the cosine similarity between the two word vectors. The cosine similarity between two vectors is calculated based on the following formulas (Equation 1) to (Equation 3). Furthermore, the more similar the two target vectors are, the closer the cosine similarity value is to 1, and the more dissimilar the two vectors are, the closer the cosine similarity value is to -1.

[0048]

number

[0049] Referring again to Figure 3, the weighting processing unit 107 assigns weights to a series of data from which feature amounts have been extracted and associated by the data analysis unit 103, for each task that is a candidate for classification, based on the similarity determination result for each of the series of data by the similarity determination unit 106. Specifically, the weighting processing unit 107 assigns weights to the series of data so that data with a higher word vector similarity to the task that is a candidate for classification is given higher priority.

[0050] An example of the processing of the weighting processing unit 107 will now be described in more detail with reference to Fig. 5, focusing on a case where the target data is moving image data corresponding to the image capture results by an imaging device. Fig. 5 is an explanatory diagram outlining an example of a case where weights are set for still images corresponding to each of a series of frames constituting a moving image, for each task that is a candidate for classification, based on the similarity determination result by the similarity determination unit 106.

[0051] As shown in the example of FIG. 5, a series of frames constituting a moving image includes frames that are highly related to the target task and frames that are less related to the task. For example, when targeting pruning work, features with a higher degree of relevance to the work are extracted from frames in which tools used in the work, such as gloves or stepladders, are captured as subjects. In other words, word vectors extracted from still images corresponding to frames with a high degree of relevance to the target work exhibit a higher degree of similarity to word vectors extracted based on work events corresponding to the work. Due to these characteristics, when frames with a high degree of relevance to the target work are used to attempt to identify the work being performed by a worker based on features extracted from still images corresponding to the frames, the work will be classified with a high probability as the target work. On the other hand, for frames in which only trees or grass are captured as subjects, the image corresponding to the frame alone may suggest a relationship to other tasks, regardless of whether the task is pruning or not. Therefore, these frames have a lower degree of relevance to pruning than frames in which gloves, ladders, or other objects are captured. In other words, word vectors extracted from still images corresponding to frames with a lower degree of relevance to the target task exhibit a lower similarity to word vectors extracted based on task events corresponding to the task. Given these characteristics, if an attempt is made to identify the task being performed by a worker using a frame with a lower degree of relevance to the target task based on features extracted from the still image corresponding to the frame, the task may be classified as a different task from the target task.

[0052] Under such circumstances, for example, suppose an attempt is made to classify the work being performed by a worker based on the average of feature amounts extracted from still images corresponding to each of the series of frames shown in FIG. 5. In the example shown in FIG. 5, as described above, the series of frames to be targeted includes frames with a low relevance to pruning work. Therefore, when feature amounts corresponding to each of the series of frames are used to classify the work being performed by the worker, the greater the number of frames with a low relevance, the lower the certainty that the work being performed by the worker is pruning work. Therefore, in this case, there is a possibility that the work being performed by the worker will be classified as a work other than pruning work.

[0053] Therefore, for each task that is a candidate for classification, the weighting processing unit 107 sets weights for the series of frames based on the results of determining the similarity between each of the series of frames and the word vector, so that frames with higher similarity are given higher priority. For example, in the example shown in Fig. 5, when it is determined whether the work being performed by the worker is classified as pruning work, the weighting processing unit 107 sets weights for a series of frames so that feature amounts corresponding to frames highly relevant to the pruning work are given higher priority. As a result, in the example shown in Fig. 5, when it is determined whether the work being performed by the worker is classified as pruning work, feature amounts extracted from still images corresponding to frames highly relevant to the pruning work are given more consideration.

[0054] Referring again to FIG. 3, the classification unit 108 classifies a series of target data, i.e., a series of data for which feature values ​​have been extracted and associated by the data analysis unit 103, according to the feature values. In this case, the classification unit 108 may divide the series of target data into groups each containing at least one piece of data, and then classify the series of data on a group-by-group basis. In this case, when grouping the series of data, the classification unit 108 may set the groups so that a predetermined number of pieces of data whose observation timings are consecutive in chronological order are assigned to a common group. Furthermore, the classification unit 108 may classify the series of data taking into account the weights assigned to the series of data by the weighting processing unit 107. Note that the method is not particularly limited as long as it is possible to classify the series of data according to the feature values. Note that hereinafter, it is assumed that the series of data is classified using a technique called clustering, so that data with more similar feature values ​​are classified into the same group.

[0055] As a specific example, the classification unit 108 may set groups of a series of frames constituting a moving image such that a predetermined number of frames whose image capture timing (i.e., observation timing) is consecutive in chronological order are assigned to a common group. As a result, the period during which the moving image was captured (in other words, the observation period) is divided into multiple periods each having a predetermined time width, and for each period, a group is set to which still images corresponding to the frames included in that period are assigned. In the following description, for convenience, each group (i.e., a group including multiple frames) set in accordance with the grouping of the series of frames will also be referred to as a frame group. Then, the classification unit 108 classifies each frame group based on the feature values ​​(e.g., word vectors) corresponding to the frames included in the frame group. At this time, the classification unit 108 may calculate the feature values ​​for each frame group based on the feature values ​​corresponding to each of the multiple frames included in the frame group, and then classify the frame groups based on the feature values ​​for each frame group. Furthermore, when calculating the feature amount for each frame group, the classification unit 108 may take into consideration the weights set for the feature amounts corresponding to the series of frames included in the frame group. As a specific example, when calculating the feature amount related to the degree of certainty that the series of frames included in the target frame group are image capture results of the implementation of pruning work by a worker, the classification unit 108 may give priority to the feature amount corresponding to the frame that has a high degree of relevance to the pruning work.

[0056] Then, the classification unit 108 outputs the classification result of the target series of data to a predetermined output destination. As a specific example, the classification unit 108 may store the classification result of the target series of data in the storage unit 110. In this way, by using the classification result of the series of data stored in the storage unit 110, it becomes possible to classify, recognize, identify, or estimate the work being performed by a predetermined worker (for example, a worker wearing the observation device 310) in the environment where the information recorded as the data was observed. Specifically, the task being performed by the worker can be classified, recognized, identified, or estimated depending on whether the characteristics of the group into which the target data is classified by clustering or the like more closely resemble the characteristics of one of the series of tasks that are candidates for classification. The target data may also include information regarding the date and time when observation was performed by the observation device 310. In this case, the date and time information included in the target data can be used to identify the date and time when the worker was performing the task (for example, the start and end times of the task). Furthermore, as described above, a series of target data may be divided into groups, and then the data may be classified for each group. In this case, by using the classification results corresponding to each group, it is possible to classify, recognize, identify, or estimate the work being performed by a specific worker. Therefore, for example, by dividing a series of observation periods into multiple partial periods and setting groups for each of the multiple partial periods, it is possible to classify, recognize, identify, or estimate the work being performed by a worker for each partial period.

[0057] It should be noted that the above-described configuration is merely an example, and the functional configuration of the information processing system 1 (particularly the functional configuration of the server device 100) is not necessarily limited to the example shown in FIG. 3. For example, a series of components of the server device 100 may be realized by multiple devices working together. As a specific example, some of the series of components of the server device 100 may be externally attached to the server device 100. As another example, the load related to the processing of at least some of the series of components of the server device 100 may be distributed to multiple devices.

[0058] An example of the functional configuration of the information processing system 1 according to this embodiment has been described above with reference to FIGS. 3 to 5, focusing particularly on the configuration of the server device 100.

[0059] <Processing> 6 and 7, an example of processing by the information processing system 1 according to this embodiment will be described, focusing particularly on processing by the server device 100. FIG. 6 is a flowchart showing an example of processing by the server device 100 according to this embodiment. In the example shown in FIG. 6, an imaging device capable of capturing moving images is applied as the observation device 310, and moving image data corresponding to the imaging results of the environment in which the worker performs work is to be analyzed by the server device 100. FIG. 7 is an explanatory diagram for describing an example of analysis processing by the server device 100 targeting image data of moving images.

[0060] In S101, the server device 100 acquires image data of a moving image corresponding to an imaging result by the observation device 310 (imaging device) from the wearable device 300 supporting the observation device 310 via a network. In S102, the server device 100 divides the image data of the moving image acquired in S101 into image data of still images corresponding to the series of frames that make up the moving image.

[0061] In S103, the server device 100 extracts, for each frame, a feature amount indicating characteristics of the surrounding situation of the worker wearing the wearable device 300 from the still image corresponding to that frame, and tags the image data of the still image with information indicating the extracted feature amount. Here, the server device 100 uses a technology called "Image to Text" to tag the image data of the still image with text information indicating an object captured as a subject in the still image (e.g., a tool used for work). For example, in the example shown in FIG. 7, for each of a series of frames constituting a moving image, information about the subject is extracted as text information from the still image corresponding to that frame, and the extracted text information about the subject is tagged as a feature amount to the image data of the still image.

[0062] In S104, the server device 100 sets a weight for the image data for each frame based on the information tagged to the image data for each frame in S103. As a specific example, the server device 100 may set a weight for the image data based on the similarity between a word vector based on the information tagged to the image data for each frame and a word vector for each task that is a candidate for classification. As a result, for example, the target image data is weighted so that the higher the degree of relevance of the features of the subject captured in the corresponding still image with the task being compared for similarity, the higher the confidence that the image data indicates the situation in which the worker is performing the task.

[0063] In S105, the server device 100 determines whether the analysis processes shown in S103 and S104 have been performed up to the final frame of the target video (that is, the video corresponding to the image data acquired in S101). If the server device 100 determines in S105 that the analysis process has not been performed up to the final frame, the process proceeds to S103. In this case, the server device 100 performs the analysis process shown in S103 and S104 on frames that have not yet been subjected to the analysis process. If the server device 100 determines in S105 that the analysis process has been executed up to the final frame, the process proceeds to S106.

[0064] In S106, the server device 100 sets frame groups by dividing a series of frames of the target video (i.e., the video corresponding to the image data acquired in S101) into groups each including at least one frame. For example, in the example shown in Fig. 7, a plurality of frame groups FG are set, each including a predetermined number of frames.

[0065] In S107, the server device 100 classifies the image data for each frame group set in S106 based on features (e.g., word vectors) extracted from image data corresponding to each of the series of frames included in the frame group. In the example shown in Fig. 6, the server device 100 classifies the image data for each frame group by so-called clustering. In addition, at this time, the server device 100 may perform clustering of the image data for the target frame group, taking into account the weights set for the image data corresponding to each frame in S104. In S108, the server device 100 outputs the classification result of the image data in S107 (e.g., the clustering result) to a predetermined output destination. As a specific example, the server device 100 may store the classification result of the image data in a predetermined storage area (e.g., the storage unit 110). As a result, by using the classification results of the series of image data, the server device 100 can classify, recognize, identify, or estimate the work being performed by the worker wearing the observation device 310 in the environment in which the information recorded as the image data was observed. Also, as described above, by classifying the image data on a frame group basis, the server device 100 can, for example, classify, recognize, identify, or estimate the work being performed by the worker during each period corresponding to a frame group.

[0066] In S109, the server device 100 determines whether the analysis processes shown in S107 and S108 have been executed up to the final frame group of the series of frame groups set in S106. If the server device 100 determines in S109 that the analysis process has not been performed up to the final frame group, the process proceeds to S107. In this case, the server device 100 performs the analysis process shown in S107 and S108 on the frame groups that have not yet been subjected to the analysis process. If the server device 100 determines in S109 that the analysis process has been performed up to the final frame group, the server device 100 ends the series of processes shown in FIG.

[0067] An example of the processing of the information processing system 1 according to this embodiment has been described above with reference to FIGS. 6 and 7, focusing particularly on the processing of the server device 100.

[0068] <Modification> Next, as a modified example of the information processing system according to this embodiment, an example will be described in which a sound collection device such as a microphone is applied as the observation device 310, and acoustic data corresponding to the sound collection results by the sound collection device is subjected to analysis processing.

[0069] First, referring to Fig. 8, an example of a method for extracting features indicating characteristics of the situation around a worker using sound data corresponding to the sound collection results by the observation device 310 (sound collection device) and using the features to classify the work being performed by the worker will be described. Fig. 8 is a diagram outlining a process for determining the similarity between the observed situation around a worker and the characteristics of the situation around the worker that would be expected if the worker were performing work that is a candidate for classification, as an example of a process for determination using features extracted from sound. In the example shown in Fig. 8, a trained model is used to extract features from the target sound, and certainty information indicating the likelihood that the target indicated by the sound is a predetermined detection target is extracted as the feature.

[0070] First, the flow shown on the left side of the example shown in Fig. 8 will be described. In the example shown in Fig. 8, the flow shown on the left side shows a processing flow related to the extraction of feature quantities from the sound collection results of sound propagating in the environment around the worker by a sound collection device. Specifically, in the example shown in Fig. 8, sound corresponding to the sound collection results is converted into a spectrogram image, and the spectrogram image is divided into predetermined periods. Then, by inputting the spectrogram image for each divided period into a trained model, certainty information indicating the likelihood that the target indicated by the sound corresponding to the spectrogram is the predetermined detection target is extracted.

[0071] Here, an overview of the spectrogram image will be explained with reference to Fig. 9. Fig. 9 shows an example of a spectrogram image converted from sound corresponding to the sound collection results by a sound collection device. In the spectrogram image shown in Fig. 9, the horizontal axis represents time and the vertical axis represents frequency. Furthermore, the brightness and color of each dot represent the strength (amplitude) of the frequency component corresponding to the position in the vertical axis direction at the time corresponding to the position in the horizontal axis direction.

[0072] Here, reference is made again to Figure 8. In the example shown in Figure 8, it is assumed that environmental sounds generated when cutting branches with a blade such as scissors during pruning work are collected. In this case, the certainty information is set such that the probability that the sound according to the sound collection results is the "sound of cutting" branches with a blade such as scissors is higher, and the probability that it is sound that is not actually collected, such as conversation, is lower.

[0073] 8, feature vectors (i.e., word vectors) defined by the probability that the observed object is each of a series of candidates defined as detection targets are extracted based on the confidence information output from the trained model. Specifically, a word vector is extracted for each of the series of candidates defined as detection targets, and a word vector corresponding to the sound corresponding to the sound collection result is extracted by taking a weighted average based on the confidence of each candidate in the confidence information for the word vector of each candidate (in other words, predicted probability).

[0074] Next, the flow shown on the right side of the example shown in Fig. 8 will be described. In the example shown in Fig. 8, the flow shown on the right side, like the example shown in Fig. 4, shows a processing flow related to the extraction of feature quantities for each task that is a candidate for classification. That is, for each task defined as a candidate for classification, a word vector is extracted based on the task event defined for that task. Note that in the example shown in Fig. 8, when extracting the word vector, information on particularly observable sounds, such as information defined as "sound effects," may be used from among the information defined as task events.

[0075] Then, the server device 100 determines the similarity between the word vector based on the confidence factor information corresponding to the sound collection results of the sound propagating in the environment around the worker by the sound collection device and the word vector extracted based on the task event for each task that is a candidate for classification. The similarity determined in this way is used to set a weight for the target data (in this modified example, the sound data corresponding to the sound collection results), as in the above-mentioned embodiment.

[0076] Next, an example of the processing of the information processing system 1 according to this modification will be described with reference to Fig. 10, focusing particularly on the processing of the server device 100. Fig. 10 is a flowchart showing an example of the processing of the server device 100 according to this modification. In the example shown in Fig. 10, a so-called video camera is used as the observation device 310, and video and audio data corresponding to the observation results of the environment in which the worker performs the work, in particular audio data, are targeted for analysis by the server device 100.

[0077] In S201, the server device 100 acquires image data of a moving image corresponding to the imaging results of the observation device 310 (video camera) via a network from the wearable device 300 supporting the observation device 310. Note that the image data includes sound data corresponding to the sound collection results of a sound collection device such as a microphone provided in the video camera. In S202, the server device 100 extracts sound data corresponding to the sound collection results by the sound collection device provided in the video camera from the image data of the moving image acquired in S201. In S203, the server device 100 converts the sound data extracted in S202 into a spectrogram image. In S204, the server device 100 divides the spectrogram image into which the acoustic data has been converted in S203 into periods of a predetermined length along a time series. Here, it is assumed that the server device 100 divides the target spectrogram image into frames (i.e., into periods corresponding to frames).

[0078] In S205, the server device 100 extracts, for each frame, features indicating the characteristics of the surrounding situation of the worker wearing the wearable device 300 from the spectrogram image corresponding to that frame, and tags the data of the spectrogram image with information indicating the extracted features.

[0079] In S206, the server device 100 sets a weight for the spectrogram image data for each frame based on the information tagged to the data in S205. As a specific example, the server device 100 may set a weight for the data based on the similarity between a word vector based on the information tagged to the spectrogram image data for each frame and a word vector for each task that is a candidate for classification. As a result, for example, the target spectrogram image data is weighted so that the higher the degree of relevance of the characteristics of the corresponding sound generation factor with the task that is being compared for similarity, the higher the confidence that the data indicates that the worker is performing the task.

[0080] In S207, the server device 100 determines whether or not the analysis processes shown in S205 and S206 have been executed up to the final frame of the series of frames corresponding to the period during which the target sound was collected. If the server device 100 determines in S207 that the analysis process has not been performed up to the final frame, the process proceeds to S205. In this case, the server device 100 performs the analysis process shown in S205 and S206 on frames that have not yet been subjected to the analysis process. If the server device 100 determines in S207 that the analysis process has been executed up to the final frame, the process proceeds to S208.

[0081] The processes in S208 to S211 are substantially the same as the processes in S106 to S109 in the example described with reference to FIG. 6, except that the target of the processes is spectrogram image data, and therefore detailed description thereof will be omitted.

[0082] In this way, the spectrogram image for each frame (in other words, the sound collected during the period corresponding to the frame) is classified according to the extracted feature. By using the classification results of the series of spectrogram image data, the server device 100 can classify, recognize, identify, or estimate the work being performed by the worker wearing the observation device 310 in the environment where the information recorded as the data was observed. Furthermore, as described above, by classifying the spectrogram image data (in other words, the sound converted into the spectrogram image) on a frame group basis, the server device 100 can, for example, classify, recognize, identify, or estimate the work being performed by the worker during each period corresponding to the frame group.

[0083] 10 shows a case where sound corresponding to the sound collection result is converted into a spectrogram image and then feature amounts are extracted from the spectrogram image, but the method is not limited as long as feature amounts indicating characteristics of the surrounding situation of the sound collection device can be extracted from the target sound. For example, feature amounts indicating characteristics of the surrounding situation of the sound collection device may be extracted by analyzing the sound data itself based on characteristics of the sound such as frequency, amplitude, phase, and distortion.

[0084] In addition, in the example shown in Figure 10, we have explained an example in which acoustic data contained in image data of a moving image based on the results of imaging by a video camera is used to classify the work being performed by a worker, but the image data of the moving image itself can also be used in a similar manner to the example shown in Figure 4.

[0085] Furthermore, the work performed by the target worker may be classified, recognized, identified, or estimated by combining the analysis results of the video image data and the analysis results of the audio data. In this case, the server device 100 may prioritize the analysis results of the video image data and the analysis results of the audio data based on predetermined conditions (e.g., conditions at the time of observation, etc.). As a specific example, when observation is performed in the evening or at night, the environment is darker than during the daytime, which may reduce the accuracy of detecting the subject from the image based on the imaging results, and ultimately reduce the accuracy of extracting feature quantities that indicate the characteristics of the situation around the worker. Therefore, under such circumstances, the server device 100 may prioritize the analysis results of the acoustic data when classifying the work being performed by the worker. As another example, in an environment with a strong influence of noise, the sound to be detected may be drowned out by the noise, reducing the accuracy of the analysis of the sound, and as a result, reducing the accuracy of the extraction of feature quantities that indicate the characteristics of the situation around the worker. Therefore, in such a situation, the server device 100 may prioritize the analysis results of image data such as moving images and still images when classifying the work being performed by the worker.

[0086] Above, with reference to Figures 8 to 10, an example of a modified example of the information processing system according to this embodiment has been described, in which a sound collection device such as a microphone is applied as the observation device 310, and acoustic data corresponding to the sound collection results by the sound collection device is subjected to analysis processing.

[0087] <Conclusion> As described above, in one embodiment of the present disclosure, an information processing device (e.g., server device 100) associates data based on observation results of a situation around a specific subject by an observation device worn by the subject with incidental information related to characteristics of the situation around the subject that corresponds to the observation results. Furthermore, for each task that is a candidate for classification, the information processing device assigns weights to a series of data associated with the incidental information so that data associated with the incidental information related to characteristics that are more similar to features based on information previously registered for the task is given higher priority. The information processing device then classifies the tasks being performed by the subject based on the weighted data for each task that is a candidate for classification. With the above configuration, when classifying the tasks being performed by a target worker, more consideration is given to feature quantities extracted from data that are more highly relevant to the task being performed by the worker, among a series of data corresponding to the observation results of the worker's surroundings. Therefore, the information processing system according to this embodiment makes it possible to achieve more accurate classification of the tasks being performed by the worker using data based on the observation results of the worker's surroundings.

[0088] It should be noted that the above-described embodiment is merely an example and does not necessarily limit the configuration or processing of the present invention, and various modifications and changes may be made without departing from the technical concept of the present invention.

[0089] For example, in the above-described embodiment and modified examples, an example of classifying, recognizing, identifying, or estimating the work being performed by a worker has been described, but the subject of analysis is not necessarily limited to so-called people such as workers, as long as it is an entity that performs the work. As a specific example, equipment (e.g., a work vehicle, etc.) that a worker uses when performing various tasks may be analyzed, and the work being performed by the equipment (in other words, the work being performed by the worker using the equipment) may be classified, recognized, identified, or estimated.

[0090] Furthermore, in the above-described embodiment and modified examples, an example of observing the situation around a target worker using an observation device worn by the target worker has been described. However, as long as it is possible to observe the situation around the target worker while the target worker is performing work, the installation location of the observation device used for the observation is not necessarily limited. As a specific example, in a situation where the environment to be observed is relatively small, the observation device (e.g., an imaging device, a sound collection device, etc.) may be installed in a position where it can capture the environment within its observation range. Furthermore, multiple observation devices may be used to observe the situation around the target worker, and multiple observation devices of different types, such as an imaging device and a sound collection device, may be used. Furthermore, in a situation where multiple observation devices are used, the devices may be installed in different positions.

[0091] Furthermore, in the above-described embodiment and modified examples, an example has been described in which the situation around a target worker is observed visually or auditorily, and the results of the observation are used to classify, recognize, identify, or estimate the work being performed by the target worker. On the other hand, as long as the situation around a target worker is observed and the results of the observation can be used to classify, recognize, identify, or estimate the work being performed by the target worker, the observation target, observation method, observation configuration, etc. are not particularly limited. In other words, by acquiring information from the five senses other than sight and hearing as observation results, the observation results may be used to classify, recognize, identify, or estimate the work being performed by the worker. As a specific example, tactile information when a worker grasps an object such as a tool to perform a task, or force information applied to the object, may be acquired as observation results, and the observation results may be used to classify, recognize, identify, or estimate the task being performed by the worker. As another example, if it is possible to detect odors around the worker as olfactory information, the olfactory information may be used as observation results to classify, recognize, identify, or estimate the task being performed by the worker. Furthermore, like the combination of visual information (images based on imaging results) and auditory information (sound based on sound collection results), it is also possible to classify, recognize, identify, or estimate the work being performed by a worker by combining and using observation results corresponding to each of multiple modalities. In this case, prioritization may be performed regarding which of the observation results corresponding to each of the multiple modalities should be used preferentially for classifying, recognizing, identifying, or estimating the work being performed by the worker, depending on predetermined conditions (e.g., observation conditions, etc.). This can be expected to further improve the accuracy of classification, recognition, identification, or estimation of the work being performed by a worker by prioritizing the modal that can observe the status of the work environment more accurately, depending on the observation conditions of the work environment.

[0092] The present invention also includes a program for realizing the functions of the above-described embodiments, and a computer-readable recording medium storing the program. [Explanation of symbols]

[0093] 1. Information Processing Systems 100 Server device 101 Communications Department 102 Input / output control unit 103 Data Analysis Department 104 Feature Extraction Unit 105 Auxiliary Processing Department 106 Similarity determination unit 107 Weighting processing unit 108 Classification Department 110 Storage section 200 Terminal Device 300 Wearable Devices 310 Observation Equipment

Claims

1. an association means for associating, as supplementary information, information relating to a detection target indicated by sound, with sound data based on a sound collection result obtained by a sound collection device attached to a predetermined subject, based on an analysis result of the sound; a weighting means for assigning weights to a series of data associated with the additional information so that data associated with the additional information relating to features having a higher similarity to features based on information previously registered for each task that is a candidate for classification is given higher priority; a classification means for classifying the tasks being performed by the subject based on the data in which the weight for each task that is a candidate for classification is set; Equipped with the weighting means, based on the result of analysis of a spectrogram image converted from sound corresponding to the sound collection result by the sound collection device, associates information about the detection target indicated by the sound as the supplementary information with the data corresponding to the sound, and sets a weight for the series of data associated with the supplementary information so that, for each task that is a candidate for classification, priority is given to the data associated as the supplementary information with certainty information that indicates a feature with a higher similarity to a feature based on one or more detection targets pre-registered for the task; The associating means inputs the spectrogram image converted from the sound based on the sound collection result by the sound collection device into a recognizer constructed based on machine learning, and associates certainty information output from the recognizer, which indicates the likelihood that the object indicated by the sound is the detection object, as the additional information, with the data corresponding to the sound. Information processing device.

2. For each task that is a candidate for classification, one or more pieces of character information related to the task are registered in advance; the associating means associates the supplementary information, which indicates a detection target indicated by sound based on a sound collection result by the sound collection device and is associated with one or more pieces of character information, with the data; the weighting means sets the weight for the series of data based on a similarity between a first feature vector based on the one or more pieces of character information registered for the task to be classified and a second feature vector based on the additional information associated with each of the series of data; The information processing device according to claim 1 .

3. the weighting means calculates a similarity between the first feature quantity vector and the second feature quantity vector based on an inner product of the first feature quantity vector and the second feature quantity vector; The information processing device according to claim 2 .

4. The classification means classifies each of the series of data associated with the incidental information according to the incidental information associated with the data and the weight set for the data, and classifies the work being performed by the entity based on the classification result of the series of data. The information processing device according to any one of claims 1 to 3.

5. For the tasks that are candidates for classification, information on factors that cause environmental sounds is registered as at least some of the detection targets; the associating means associates, based on an analysis result of the sound based on the sound collection result by the sound collection device, information on a cause of the sound as the additional information with the data corresponding to the sound; The weighting means sets a weight for the series of data associated with the incidental information so that, for each task that is a candidate for classification, a feature based on information on the factors that cause the environmental sound that has been registered in advance for the task is given higher priority to the data associated with information on the factors that cause the sound that shows a higher similarity as the incidental information. The information processing device according to any one of claims 1 to 4.

6. An information processing method executed by an information processing device, an associating step of associating information about a detection target indicated by sound as supplementary information with sound data based on a sound collection result by a sound collection device attached to a predetermined subject, based on an analysis result of the sound; a weighting step of setting weights for a series of data associated with the additional information so that data associated with the additional information relating to features having a higher similarity to features based on information previously registered for each task that is a candidate for classification is given higher priority; a classification step of classifying the tasks being performed by the subject based on the data in which the weights for each task that is a candidate for classification are set; Including, The weighting step associates, as the supplementary information, information relating to the detection target indicated by the sound based on the analysis result of an image of a spectrogram converted from sound corresponding to the sound collection result by the sound collection device, with the data corresponding to the sound, and sets weights for the series of data associated with the supplementary information so that, for each task that is a candidate for classification, priority is given to the data associated as the supplementary information with certainty information indicating a feature that is more similar to a feature based on one or more detection targets pre-registered for the task, The associating step inputs the spectrogram image converted from the sound based on the sound collection result by the sound collection device to a recognizer constructed based on machine learning, and associates certainty information output from the recognizer, which indicates the likelihood that the object indicated by the sound is the detection object, as the additional information, with the data corresponding to the sound. Information processing methods.

7. On the computer, an associating step of associating information about a detection target indicated by sound as supplementary information with sound data based on a sound collection result by a sound collection device attached to a predetermined subject, based on an analysis result of the sound; a weighting step of setting weights for a series of data associated with the additional information so that data associated with the additional information relating to features having a higher similarity to features based on information previously registered for each task that is a candidate for classification is given higher priority; a classification step of classifying the tasks being performed by the subject based on the data in which the weights for each task that is a candidate for classification are set; Execute The weighting step associates, as the supplementary information, information relating to the detection target indicated by the sound based on the analysis result of an image of a spectrogram converted from sound corresponding to the sound collection result by the sound collection device, with the data corresponding to the sound, and sets weights for the series of data associated with the supplementary information so that, for each task that is a candidate for classification, priority is given to the data associated as the supplementary information with certainty information indicating a feature that is more similar to a feature based on one or more detection targets pre-registered for the task, The associating step inputs the spectrogram image converted from the sound based on the sound collection result by the sound collection device to a recognizer constructed based on machine learning, and associates certainty information output from the recognizer, which indicates the likelihood that the object indicated by the sound is the detection object, as the additional information, with the data corresponding to the sound. program.

Citation Information

Patent Citations

  • Information processing device, electronic instrument, information processing method and program

    JP2013254372A

  • Event series extraction apparatus, event series extraction method, and event extraction program

    JP2019046304A

  • System and method for determining work by work vehicle, and learned model manufacturing method

    JP2020004096A

  • Notification devices and wearable devices

    JP3233390U