Information processing method, information processing device, and information processing program

The method addresses excessive data storage and worker burden by initiating video and audio recording at a work site based on remote support triggers, ensuring efficient and selective data capture.

JP2026071426APending Publication Date: 2026-04-30PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2024-01-30
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

The conventional methods of recording all moving images and voices at a work site result in excessive data storage requirements, increasing memory needs and worker burden, and fail to allow selective recording of necessary data.

Method used

An information processing method that initiates video and audio recording at a work site based on triggers such as remote support worker operations, location, object identification, or predefined keywords, reducing the need for on-site worker intervention.

Benefits of technology

Reduces data storage needs and worker burden by selectively recording only necessary video and audio data, ensuring efficient use of memory resources and minimizing disruption during work operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026071426000001_ABST
    Figure 2026071426000001_ABST
Patent Text Reader

Abstract

This technology reduces the amount of data stored in memory and alleviates the burden on workers. [Solution] The server receives video footage and audio recordings taken at the work site from a work terminal used by a worker at the work site, and starts recording the video footage and audio recordings to memory when a support worker assists the worker with starting the recording operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technique for recording moving images captured at a work site and voices collected at the work site on a server.

Background Art

[0002] For example, Patent Document 1 discloses that an intermediary terminal device automatically records support request information consisting of an image signal and a voice signal passing through the intermediary terminal device and support information consisting of a voice signal.

[0003] Also, for example, Patent Document 2 discloses a client terminal including data acquisition means for acquiring data from input / output means or / and communication means, application request means for requesting an application necessary for the client terminal to function, application storage means for storing an application received from a server, and security means for deleting data acquired by the data acquisition means or / and an application stored in the application storage means when the operation of the client terminal ends.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the above conventional technology, when all moving images captured at the work site and all voices collected at the work site are recorded in the memory of the server, the amount of data to be recorded in the memory may increase, and further improvement has been required.

[0006] This disclosure was made to solve the above-mentioned problems and aims to provide a technology that can reduce the amount of data recorded in memory and alleviate the burden on workers. [Means for solving the problem]

[0007] The information processing method relating to this disclosure is an information processing method performed by a computer, and includes receiving video footage and first audio recordings taken at a work site from a work terminal used by a worker at the work site, and starting to record the video footage and first audio recordings into memory, triggered by a recording start operation by a support worker assisting the worker with their work. [Effects of the Invention]

[0008] According to this disclosure, the amount of data to be recorded in memory can be reduced, and the burden on the worker can be lessened. [Brief explanation of the drawing]

[0009] [Figure 1] This diagram shows the configuration of the work support system according to this embodiment 1. [Figure 2] This is a flowchart illustrating the work support processing performed by the work terminal in Embodiment 1 of this disclosure. [Figure 3] This is a flowchart illustrating the server-based work support processing in Embodiment 1 of this disclosure. [Figure 4] This is a flowchart illustrating the work support processing by the support terminal in Embodiment 1 of this disclosure. [Figure 5] This figure shows an example of a screen displayed on the display unit of the support terminal in this embodiment 1. [Figure 6] This diagram shows the configuration of the work support system according to this second embodiment. [Figure 7]This is a flowchart illustrating the work support processing performed by the work terminal in Embodiment 2 of this disclosure. [Figure 8] This is a flowchart illustrating the server-based work support processing in Embodiment 2 of this disclosure. [Figure 9] This is a flowchart illustrating the work support processing by the support terminal in Embodiment 2 of this disclosure. [Figure 10] This diagram shows the configuration of the work support system according to this third embodiment. [Figure 11] This is a flowchart illustrating the work support processing by the work terminal in Embodiment 3 of this disclosure. [Figure 12] This is a flowchart illustrating the server-based work support processing in Embodiment 3 of this disclosure. [Figure 13] This is a flowchart illustrating the work support processing by the support terminal in Embodiment 3 of this disclosure. [Figure 14] This diagram shows the configuration of the work support system according to this fourth embodiment. [Figure 15] This is a flowchart illustrating the server-based work support processing in Embodiment 4 of this disclosure. [Figure 16] This diagram shows the configuration of the work support system according to this embodiment 5. [Figure 17] This is a flowchart illustrating the server-based work support processing in Embodiment 5 of this disclosure. [Figure 18] This diagram shows the configuration of the work support system according to this embodiment 6. [Figure 19] This is a flowchart illustrating the work support processing performed by the work terminal in Embodiment 6 of this disclosure. [Figure 20] This is a flowchart illustrating the server-based work support processing in Embodiment 6 of this disclosure. [Figure 21] This diagram shows the configuration of the work support system according to this embodiment 7. [Figure 22] A flowchart for explaining the work support process by the server in Embodiment 7 of the present disclosure. [Figure 23] A diagram showing the configuration of the work support system according to Embodiment 8 of the present disclosure. [Figure 24] A flowchart for explaining the work support process by the work terminal in Embodiment 8 of the present disclosure. [Figure 25] A flowchart for explaining the work support process by the server in Embodiment 8 of the present disclosure. [Figure 26] A flowchart for explaining the work support process by the support terminal in Embodiment 8 of the present disclosure. [Figure 27] A diagram showing an example of the screen displayed on the display unit of the support terminal in Embodiment 8. [Figure 28] A diagram showing an example of the screen displayed on the display unit of the work terminal in Embodiment 8. [Figure 29] A diagram showing the configuration of the work support system according to Embodiment 9 of the present disclosure. [Figure 30] A flowchart for explaining the work support process by the work terminal in Embodiment 9 of the present disclosure. [Figure 31] A flowchart for explaining the work support process by the server in Embodiment 9 of the present disclosure. [Figure 32] A diagram showing the configuration of the work support system according to Embodiment 10 of the present disclosure. [Figure 33] A flowchart for explaining the work support process by the server in Embodiment 10 of the present disclosure. [Figure 34] A diagram showing an example of the screen displayed on the display unit of the support terminal in Embodiments 1 to 10.

Embodiments for Carrying Out the Invention

[0010] (Findings on which the present disclosure is based) In manufacturing or construction industries, there are workers who perform tasks on-site and support staff who assist them. When a worker needs assistance with their work, they request it from a support staff member. The support staff member may be on-site or located in a remote location. Video and audio recordings of the worker receiving support from the support staff member are stored in the server's memory, allowing the worker to review the recorded footage later.

[0011] In the above-mentioned Patent Document 1, the intermediary terminal device automatically records support request information consisting of image signals and audio signals, and support information consisting of audio signals, which pass through the intermediary terminal device.

[0012] However, if all video footage and audio recordings taken at the work site are stored in the server's memory, the amount of data stored in memory may increase. In this case, the memory's storage capacity would need to be increased, incurring additional costs. Furthermore, if workers need to initiate recording during their work to record only the necessary data, this could place a heavy burden on them and disrupt their work.

[0013] Furthermore, in the aforementioned Patent Document 2, the data acquired and / or the application is deleted when the operation of the client terminal is completed. Therefore, since the video and audio of the worker receiving assistance from the supporter are not recorded in memory, the worker was unable to view the recorded video and audio later.

[0014] To address the above challenges, the following technologies are disclosed.

[0015] (1) An information processing method relating to one aspect of the present disclosure is an information processing method performed by a computer, which includes receiving a video image taken at a work site and a first audio recording collected at the work site from a work terminal used by a worker at the work site, and starting to record the video image and the first audio recording into memory, triggered by a recording start operation by a support worker assisting the worker with the work performed.

[0016] In this configuration, recording of video and audio received from a work terminal used by a worker at the work site is initiated when a support worker initiates recording. Therefore, only the video and audio deemed necessary by the support worker are recorded in memory, thus reducing the amount of data recorded in memory. Furthermore, since recording of video and audio begins as soon as the support worker initiates recording, the worker does not need to initiate recording during work, thus reducing the burden on the worker.

[0017] (2) The information processing method described in (1) above further includes receiving a recording start signal from a support terminal used by the supporter, which instructs the start of recording based on an input operation by the supporter located in a remote location, and the start of recording may include starting the recording of the video and the first audio to the memory, triggered by the receipt of the recording start signal.

[0018] With this configuration, when an input operation is performed by a supporter located remotely, recording of the video and the first audio to memory begins. Therefore, the operator does not need to perform any operation to start recording, and the video and the first audio can be recorded to memory based on the input operation performed by the supporter located remotely.

[0019] (3) The information processing method described in (1) above further includes receiving area designation information indicating a predetermined area at the work site designated by the supporter who is in a remote location from a support terminal used by the supporter, and further includes receiving location information indicating the current location of the work terminal from the work terminal, wherein the start of recording may include starting the recording of the video and the first audio to the memory when the location of the work terminal indicated by the location information enters the predetermined area indicated by the area designation information.

[0020] With this configuration, when the work terminal enters a predetermined area at the work site designated by a remote supporter, recording of video and audio to memory begins. Therefore, the worker does not need to perform any recording start operation; the remote supporter can record video and audio to memory by specifying a predetermined area at the work site.

[0021] (4) The information processing method described in (1) above further includes receiving identification information contained in a wireless signal transmitted from a work object at the work site designated by the supporter in a remote location from a support terminal used by the supporter, and further includes receiving signal information from the work terminal, which includes the identification information contained in the wireless signal received by the work terminal and the radio wave strength of the wireless signal measured by the work terminal, and the start of recording may include starting the recording of the video and the first audio to the memory, triggered by the radio wave strength of the wireless signal received from the work terminal, which includes the identification information received from the support terminal, being greater than or equal to a threshold.

[0022] In this configuration, when the work terminal approaches the work object at the work site designated by a remote supporter, recording of video and audio to memory begins. Therefore, the worker does not need to perform any operation to start recording; the remote supporter can record video and audio to memory by specifying the identification information of the work object at the work site.

[0023] (5) The information processing method described in (1) above further includes receiving a second sound from the support terminal used by the supporter in a remote location, and the start of recording may include starting the recording of the video and the first sound to the memory when a predetermined keyword stored in advance is included in the second sound.

[0024] In this configuration, when a remote supporter utters a predetermined keyword that has been stored in advance, the recording of the video and first audio to memory begins. Therefore, the operator does not need to perform any operation to start recording; the video and first audio can be recorded to memory simply by the remote supporter uttering the predetermined keyword.

[0025] (6) In the information processing method described in (1) above, the start of recording may include detecting a speech segment spoken by the supporter in the second audio, and triggering the recording of the video and the first audio to the memory when the detected speech segment contains a predetermined keyword that has been stored in advance in the second audio.

[0026] In this configuration, the speech segment spoken by the supporter in the second audio is detected, and if a predetermined keyword, which is stored in advance, is found to be included in the second audio within the detected speech segment, recording of the video and the first audio to memory begins. Therefore, even if the second audio contains noise, it is possible to determine with high accuracy whether the predetermined keyword is included in the second audio.

[0027] (7) In the information processing method described in (1) above, the start of recording may include recognizing the actions of the support worker at the work site from the received video, and triggering the recording of the video and the first audio to the memory when the recognized action is a predetermined action.

[0028] In this configuration, when a support person at the work site performs a predetermined action, recording of the video and first audio to memory begins. Therefore, the worker does not need to perform any operation to start recording; the video and first audio can be recorded to memory simply by the support person at the work site performing the predetermined action.

[0029] (8) In the information processing method described in (1) above, the start of recording may include starting the recording of the video and the first audio to the memory when a predetermined keyword stored in advance is included in the first audio, which includes the voice of the supporter at the work site.

[0030] In this configuration, when a support worker at the work site utters a predetermined keyword that has been stored in advance, the recording of the video and first audio to memory begins. Therefore, the worker does not need to perform any operation to start recording; the video and first audio can be recorded to memory simply by the support worker at the work site uttering the predetermined keyword.

[0031] (9) The information processing method described in (1) above further includes transmitting the received video and audio to a support terminal used by the supporter located at a remote location, and receiving a still image extracted by the supporter from the video displayed on the display unit of the support terminal, wherein the start of recording may include starting to record the video and audio to the memory when the still image is received.

[0032] In this configuration, when a remote supporter extracts still images from the video displayed on the support terminal's display unit to assist with the work, the recording of the video and the first audio to memory begins. Therefore, the worker does not need to perform any recording start operation; the remote supporter can record the video and the first audio to memory by extracting still images from the video.

[0033] (10) The information processing method described in (9) above further includes receiving a second sound from the support terminal, and the start of recording may include starting to record the moving image, the first sound, the second sound, and the still image to the memory, triggered by the reception of the still image.

[0034] With this configuration, not only the video and first audio, but also the second audio, which includes the voice of a remote supporter, and still images extracted from the video by the supporter are recorded in memory. Therefore, even after the supporter has finished assisting with the actual work, the worker can use the video, first audio, second audio, and still images recorded in memory to assist with the work.

[0035] (11) The information processing method described in (1) above further includes transmitting the received video and audio to a support terminal used by the supporter located at a remote location, and receiving a still image extracted by the supporter from the video displayed on the display unit of the support terminal, wherein the reception of the still image includes receiving the still image from the support terminal on which instruction information input by the supporter using the support terminal is superimposed, and the start of recording may include starting to record the video and audio to the memory, triggered by the superimposition of the instruction information on the still image.

[0036] In this configuration, a remote supporter extracts still images from the video displayed on the support terminal's display unit to assist with the work, and when instruction information is superimposed on the still images, recording of the video and the first audio to memory begins. Therefore, the worker does not need to perform any operation to start recording; the remote supporter can extract still images from the video and superimpose instruction information on the still images to record the video and the first audio to memory.

[0037] (12) The information processing method described in (1) above further includes receiving mode information from the work terminal indicating whether a first mode in which the worker uses the work terminal or a second mode in which the supporter uses the work terminal has been selected, and the start of recording may include starting the recording of the video and the first audio to the memory when the received mode information indicates the second mode and a predetermined keyword stored in advance is included in the first audio.

[0038] In this configuration, when a support worker assisting with work using a work terminal at the work site utters a predetermined keyword that has been stored in advance, the recording of video and audio to memory begins. Therefore, the worker does not need to perform any operation to start recording; the video and audio to memory can be recorded simply by the support worker assisting with work using a work terminal at the work site uttering a predetermined keyword.

[0039] (13) The information processing method described in (1) above further includes receiving mode information from the work terminal indicating whether a first mode in which the worker uses the work terminal or a second mode in which the supporter uses the work terminal has been selected, and the start of recording may include starting the recording of the video and the first audio to the memory when the received mode information indicates the second mode and the same object is continuously captured in a predetermined area of ​​the video for a predetermined time or longer, triggered by this.

[0040] When a support worker using a work terminal at a work site gazes at the work object, the same object will appear continuously in a predetermined area of ​​the video for a predetermined time or longer. Therefore, when a support worker uses a work terminal at the work site and the same object appears continuously in a predetermined area of ​​the video for a predetermined time or longer, recording of the video and the first audio to memory begins. Consequently, the worker does not need to perform any operation to start recording; the video and the first audio can be recorded to memory simply by a support worker using a work terminal at the work site gazing at the work object.

[0041] Furthermore, this disclosure can be implemented not only as an information processing method that performs the characteristic processing described above, but also as an information processing device having a characteristic configuration corresponding to the characteristic processing performed by the information processing method. It can also be implemented as a computer program that causes a computer to execute the characteristic processing included in such an information processing method. Therefore, the same effects as the above-described information processing method can be achieved in the following other embodiments.

[0042] (14) An information processing device according to another aspect of the present disclosure comprises a communication unit, a control unit, and a memory, wherein the communication unit receives video footage taken at a work site and first audio recordings collected at the work site from a work terminal used by a worker at the work site, and the control unit starts recording the video footage and the first audio recordings to the memory triggered by a recording start operation by a support worker assisting the worker with their work.

[0043] (15) Information processing programs according to other embodiments of the present disclosure receive video footage and first audio recordings taken at a work site from a work terminal used by a worker at the work site, and cause a computer to start recording the video footage and first audio recordings to memory, triggered by a recording start operation by a support worker assisting the worker with their work.

[0044] (16) A non-temporary computer-readable recording medium according to another aspect of the present disclosure records an information processing program which receives video footage captured at a work site and first audio recordings collected at the work site from a work terminal used by a worker at the work site, and causes a computer to start recording the video footage and first audio recordings to memory triggered by a recording start operation by a support worker assisting the worker with the work being performed.

[0045] Embodiments of this disclosure will be described below with reference to the attached drawings. Note that the embodiments described below are all specific examples of this disclosure. The numerical values, shapes, components, steps, and order of steps shown in the following embodiments are examples only and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in the independent claim representing the highest-level concept will be described as optional components. Also, in all embodiments, the contents of each can be combined.

[0046] (Embodiment 1) Figure 1 is a diagram showing the configuration of the work support system 10 according to this embodiment 1.

[0047] The work support system 10 shown in Figure 1 comprises a work terminal 1, a server 2, and a support terminal 3.

[0048] In Embodiment 1, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0049] The work terminal 1 is, for example, a wearable device attached to the worker's head. The worker performs their work at the work site while wearing the work terminal 1. The work terminal 1 may also be, for example, a smartphone or a tablet computer.

[0050] The work terminal 1 comprises, for example, a computer system including a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. The work terminal 1 may be implemented, for example, by a hardware implementation using a processing circuit, by the execution of a software program held in memory by the processing circuit or distributed from an external server, or by a combination of these hardware and software implementations.

[0051] The work terminal 1 is connected to the server 2 via network 4, enabling communication between them. Network 4 is, for example, the internet.

[0052] The work terminal 1 includes a communication unit 11, a control unit 12, a memory 13, an input unit 14, a camera 15, a microphone 16, and a speaker 17.

[0053] The control unit 12 controls the entire work terminal 1. The control unit 12 controls the operation of the communication unit 11, memory 13, input unit 14, camera 15, microphone 16, and speaker 17.

[0054] Memory 13 is a storage device capable of storing various types of information, such as RAM (Random Access Memory), SSD (Solid State Drive), or flash memory.

[0055] Camera 15 acquires video footage by photographing the work site. If the work terminal 1 is a wearable device attached to the worker's head, the video footage is from the worker's perspective.

[0056] Microphone 16 collects primary audio at the work site.

[0057] The communication unit 11 transmits video footage captured by the camera 15 and first audio collected by the microphone 16 to the server 2. The communication unit 11 also receives second audio from the server 2, which is from the area surrounding the support terminal 3 used by a support worker in a remote location.

[0058] The input unit 14 accepts various input operations from the operator. The input unit 14 includes a first start button for starting shooting with the camera 15 and starting the collection of first sound with the microphone 16. The input unit 14 also includes a first stop button for ending shooting with the camera 15 and ending the collection of first sound with the microphone 16. When the operator presses the first start button, the camera 15 starts shooting and the microphone 16 starts collecting first sound. When the operator presses the first stop button, the camera 15 stops shooting and the microphone 16 stops collecting first sound.

[0059] The input unit 14 also includes a second start button for initiating the transmission of video and audio to the server 2. The input unit 14 also includes a second end button for ending the transmission of video and audio to the server 2. When the operator presses the second start button, the communication unit 11 starts transmitting video and first audio to the server 2. When the operator presses the second end button, the communication unit 11 ends the transmission of video and first audio to the server 2.

[0060] Speaker 17 outputs the second audio received by the communication unit 11 to the outside. The second audio includes the voice of a support worker, allowing the worker to perform their work while listening to the support worker's voice output from speaker 17.

[0061] Server 2 comprises, for example, a computer system including a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as internal memory or accessible external memory that stores the control program. Server 2 may be implemented, for example, by a hardware implementation using a processing circuit, by the execution of a software program held in memory by the processing circuit or distributed from an external server, or by a combination of these hardware and software implementations.

[0062] Server 2 is connected to work terminal 1 and support terminal 3 via network 4, enabling them to communicate with each other.

[0063] Server 2 comprises a communication unit 21, a control unit 22, and a memory 23. Server 2 is an example of an information processing device.

[0064] The communication unit 21 receives video footage and audio recordings from the work site, as well as audio recordings from the work terminal 1 used by a worker at the work site. The communication unit 21 also receives audio recordings from the support terminal 3 used by a support worker at a remote location. The communication unit 21 transmits the video footage and audio recordings received from the work terminal 1 to the support terminal 3. The communication unit 21 also transmits the audio recordings received from the support terminal 3 to the work terminal 1.

[0065] The control unit 22 controls the entire server 2. The control unit 22 controls the operation of the communication unit 21 and the memory 23. The control unit 22 starts recording the video and first audio received by the communication unit 21 into the memory 23 when an assistant who supports the worker starts recording is triggered by an operation to start recording. The control unit 22 also stops recording the video and first audio received by the communication unit 21 into the memory 23 when an assistant stops recording is triggered by an operation to end recording.

[0066] Furthermore, the control unit 22 may record not only the video and first audio from the work terminal 1 to the memory 23, but also the video and first audio from the work terminal 1 and the second audio from the support terminal 3 to the memory 23. That is, the control unit 22 may start recording the video, first audio, and second audio received by the communication unit 21 to the memory 23 when the supporter initiates a recording start operation. Also, the control unit 22 may stop recording the video, first audio, and second audio received by the communication unit 21 to the memory 23 when the supporter initiates a recording end operation.

[0067] Memory 23 is a storage device capable of storing various types of information, such as RAM, HDD (Hard Disk Drive), SSD, or flash memory. Memory 23 non-temporarily records video and audio from the work terminal 1.

[0068] Furthermore, memory 23 may not only non-temporarily record the video and first audio from work terminal 1, but may also non-temporarily record the video and first audio from work terminal 1 and the second audio from support terminal 3. In other words, memory 23 may non-temporarily record the video, first audio, and second audio received by communication unit 21. In this case, memory 23 records the video, first audio, and second audio in a single file.

[0069] Furthermore, the communication unit 21 receives a recording start signal from the support terminal 3 used by the supporter, instructing the start of recording based on input operations by the supporter located remotely. The control unit 22, triggered by the communication unit 21 receiving the recording start signal, begins recording the video and first audio to the memory 23.

[0070] Furthermore, the communication unit 21 receives a recording termination signal from the support terminal 3 used by the supporter, based on input operations performed by the supporter in a remote location, instructing the end of recording. The control unit 22 terminates the recording of the video and first audio to the memory 23, triggered by the communication unit 21 receiving the recording termination signal.

[0071] Support terminal 3 is, for example, a personal computer, a smartphone, or a tablet computer.

[0072] The support terminal 3 comprises, for example, a computer system including a control program, a processing circuit such as a processor or logic circuit that executes the control program, and a recording device such as an internal memory or an accessible external memory that stores the control program. The support terminal 3 may be implemented, for example, by hardware implementation using a processing circuit, by execution of a software program held in memory by the processing circuit or distributed from an external server, or by a combination of these hardware and software implementations.

[0073] Support terminal 3 is connected to server 2 via network 4, enabling them to communicate with each other.

[0074] The support terminal 3 comprises a communication unit 31, a control unit 32, a memory 33, a display unit 34, a speaker 35, a microphone 36, and an input unit 37.

[0075] The microphone 36 collects second audio from the surrounding area of ​​the support terminal 3.

[0076] The communication unit 31 receives video footage and first audio recordings from the server 2 at the work site. The communication unit 31 also transmits second audio recordings from the surrounding area of ​​the support terminal 3, collected by the microphone 36, to the server 2.

[0077] Furthermore, the communication unit 31 sends a recording start signal to the server 2 to instruct the start of recording based on the input operation by the supporter. Also, the communication unit 21 sends a recording end signal to the server 2 to instruct the end of recording based on the input operation by the supporter.

[0078] The control unit 32 controls the entire support terminal 3. The control unit 32 controls the operation of the communication unit 31, memory 33, display unit 34, speaker 35, microphone 36, and input unit 37.

[0079] Memory 33 is a storage device capable of storing various types of information, such as RAM, HDD, SSD, or flash memory.

[0080] The display unit 34 is, for example, a liquid crystal display and displays various information. The display unit 34 displays video footage of the work site received by the communication unit 31. The video footage displayed on the display unit 34 is real-time video footage. By viewing the video footage displayed on the display unit 34, the support staff can confirm the work being done by the workers at the work site.

[0081] Speaker 35 outputs the first audio, which is collected at the work site and received by the communication unit 31, to the outside. The first audio output from speaker 35 is audio collected in real time. Supporters can assist the worker while listening to the worker's voice output from speaker 35.

[0082] The input unit 37 is, for example, a keyboard, mouse, or touch panel. The input unit 37 accepts various input operations from the supporter. The input unit 37 includes a recording start button for starting the recording of video and first audio to the server 2. The recording start button may be a button that is physically pressed by the supporter, or it may be a button that is displayed on the display unit 34 and clicked with a mouse. When the recording start button is pressed by the supporter, the communication unit 31 sends a recording start signal to the server 2 instructing it to start recording.

[0083] The input unit 37 also includes a recording stop button for ending the recording of the video and first audio to the server 2. The recording stop button may be a button that is physically pressed by the supporter, or it may be a button that is displayed on the display unit 34 and clicked with a mouse. When the recording stop button is pressed by the supporter, the communication unit 31 sends a recording stop signal to the server 2 to instruct the end of recording.

[0084] Before initiating communication between the work terminal 1, server 2, and support terminal 3, each generates a communication ID and sends the generated communication ID to the other. The work terminal 1, server 2, and support terminal 3 use the communication ID to transmit and receive video, first audio, and second audio. The communication ID is used to identify the video, first audio, and second audio.

[0085] Next, the work support processing performed by the work terminal 1, server 2, and support terminal 3 in Embodiment 1 of this disclosure will be described.

[0086] Figure 2 is a flowchart illustrating the work support processing performed by the work terminal 1 in Embodiment 1 of this disclosure.

[0087] First, in step S1, the camera 15 acquires a video image by photographing the work site. At this time, the input unit 14 receives an input operation from the worker to start acquiring the video image and the first audio.

[0088] Next, in step S2, the microphone 16 acquires a first sound at the work site.

[0089] Next, in step S3, the communication unit 11 transmits the video footage acquired by the camera 15 and the first audio recording acquired by the microphone 16 to the server 2. At this time, the input unit 14 accepts an input operation from the operator to start transmitting the video footage and the first audio recording. The communication unit 11 also transmits the video footage and the first audio recording to the server 2, with the support terminal 3 as the destination. As a result, the video footage and the first audio recording are transmitted to the support terminal 3 via the server 2.

[0090] Next, in step S4, the communication unit 11 receives the second audio signal from the support terminal 3 transmitted by the server 2.

[0091] Next, in step S5, the speaker 17 outputs the second audio signal received by the communication unit 11 to the outside.

[0092] Next, in step S6, the control unit 12 determines whether or not to terminate the transmission of the video and first audio. At this time, the input unit 14 receives an input operation from the operator to terminate the transmission of the video and first audio. If an input operation to terminate the transmission of the video and first audio is received, the control unit 12 determines to terminate the transmission of the video and first audio. If an input operation to terminate the transmission of the video and first audio is not received, the control unit 12 determines not to terminate the transmission of the video and first audio.

[0093] If it is determined at this point that the transmission of the video and first audio should be terminated (YES in step S6), the work support process ends. At this time, the communication unit 11 terminates the transmission of the video and first audio. After the transmission of the video and first audio is terminated, the input unit 14 accepts input from the operator to terminate the acquisition of the video and first audio.

[0094] On the other hand, if it is determined that the transmission of the video and the first audio should not be terminated (NO in step S6), the process returns to step S1.

[0095] Figure 3 is a flowchart illustrating the work support processing performed by Server 2 in Embodiment 1 of this disclosure.

[0096] First, in step S11, the communication unit 21 receives the video and audio transmitted by the work terminal 1.

[0097] Next, in step S12, the communication unit 21 transmits the received video and first audio to the support terminal 3.

[0098] Next, in step S13, the communication unit 21 receives the second audio transmitted by the support terminal 3.

[0099] Next, in step S14, the communication unit 21 transmits the received second audio to the work terminal 1.

[0100] Next, in step S15, the control unit 22 determines whether or not a recording start signal instructing the start of recording has been received by the communication unit 21 based on an input operation by a supporter located remotely.

[0101] If it is determined that a recording start signal has been received (YES in step S15), in step S16, the control unit 22 starts recording the video, first audio, and second audio received by the communication unit 21 into the memory 23. Then, the process returns to step S11. From this point onward, the video, first audio, and second audio received by the communication unit 21 are recorded into the memory 23.

[0102] On the other hand, if it is determined that the recording start signal has not been received (NO in step S15), in step S17, the control unit 22 determines whether or not the communication unit 21 has received a recording end signal instructing the end of recording based on an input operation by a supporter in a remote location.

[0103] If it is determined that a recording end signal has been received (YES in step S17), in step S18, the control unit 22 terminates the recording of the video, first audio, and second audio received by the communication unit 21 to the memory 23. Then, the process returns to step S11. As a result, the video, first audio, and second audio received by the communication unit 21 from the time the recording start signal is received until the time the recording end signal is received are recorded in the memory 23.

[0104] On the other hand, if it is determined that no recording end signal has been received (NO in step S17), the process returns to step S11.

[0105] Figure 4 is a flowchart illustrating the work support processing by the support terminal 3 in Embodiment 1 of this disclosure.

[0106] First, in step S21, the communication unit 31 receives the video and audio transmitted by the server 2.

[0107] Next, in step S22, the display unit 34 displays the video received by the communication unit 31.

[0108] Next, in step S23, the speaker 35 outputs the first audio signal received by the communication unit 31 to the outside.

[0109] Next, in step S24, the microphone 36 acquires a second sound from the surroundings of the support terminal 3.

[0110] Next, in step S25, the communication unit 31 transmits the second audio acquired by the microphone 36 to the server 2. At this time, the communication unit 31 transmits the second audio to the server 2 with the work terminal 1 as the destination. As a result, the second audio is transmitted to the work terminal 1 via the server 2.

[0111] Next, in step S26, the control unit 32 determines whether or not the recording start button on the input unit 37 has been pressed.

[0112] If it is determined that the record start button has been pressed (YES in step S26), then in step S27, the communication unit 31 sends a record start signal to the server 2 to instruct the start of recording. After that, the process returns to step S21.

[0113] On the other hand, if it is determined that the recording start button has not been pressed (NO in step S26), in step S28, the control unit 32 determines whether or not the recording end button on the input unit 37 has been pressed.

[0114] If it is determined that the recording stop button has been pressed (YES in step S28), then in step S29, the communication unit 31 sends a recording stop signal to the server 2 to instruct the end of recording. After that, the process returns to step S21.

[0115] On the other hand, if it is determined that the recording stop button has not been pressed (NO in step S28), the process returns to step S21.

[0116] In this way, the recording of video and audio received from the work terminal 1 used by the worker at the work site to memory 23 is initiated by the support worker's operation to start recording. Therefore, only the video and audio that the support worker deems necessary are recorded to memory 23, thus reducing the amount of data recorded to memory 23. In addition, since recording of video and audio to memory 23 begins when the support worker initiates recording, the worker does not need to initiate recording during work, thus reducing the burden on the worker.

[0117] Furthermore, when an input operation is performed by a supporter located remotely, recording of the video and the first audio to memory 23 begins. Therefore, the operator does not need to perform an operation to start recording, and the video and the first audio can be recorded to memory 23 based on the input operation performed by the supporter located remotely.

[0118] Figure 5 shows an example of a screen displayed on the display unit 34 of the support terminal 3 in this embodiment 1.

[0119] The display unit 34 displays a video 341 of the work site, a recording start button 342, and a recording stop button 343. When the assistant operates the mouse, the pointer displayed on the display unit 34 moves over the recording start button 342, and when the assistant clicks the mouse button, a recording start signal is sent to the server 2. As a result, the server 2 starts recording the video, the first audio, and the second audio.

[0120] Furthermore, while the video, first audio, and second audio are being recorded, the supporter moves the pointer displayed on the display unit 34 over the recording end button 343 by operating the mouse. When the supporter clicks the mouse button, a recording end signal is sent to the server 2. As a result, the server 2 terminates the recording of the video, first audio, and second audio.

[0121] In this embodiment 1, the control unit 22 starts recording the video and first audio to the memory 23 from the time the supporter initiates recording. However, this disclosure is not limited to this, and the video and first audio may be recorded to the memory 23 from a predetermined time before the time the supporter initiates recording. In this case, the memory 23 includes a buffer area for temporarily recording the received video and first audio. The control unit 22 temporarily records the received video and first audio in the buffer area. The control unit 22 may read the video and first audio from the buffer area up to a predetermined time before the time the recording start signal is received and record them to the memory 23, and may also record the video and first audio from the time the recording start signal is received onward to the memory 23.

[0122] Furthermore, although the control unit 22 terminates recording of the video and first audio to the memory 23 when the supporter performs the recording termination operation, this disclosure is not limited thereto, and the video and first audio may be recorded to the memory 23 from the time the supporter performs the recording termination operation until a predetermined time has elapsed. The control unit 22 may also record the video and first audio to the memory 23 from the time the recording termination signal is received until a predetermined time has elapsed.

[0123] (Embodiment 2) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 2, recording of the video and first audio to memory is initiated when the work terminal enters a predetermined area at the work site designated by a remote supporter.

[0124] Figure 6 shows the configuration of the work support system 10A according to this second embodiment.

[0125] The work support system 10A shown in Figure 6 comprises a work terminal 1A, a server 2A, and a support terminal 3A. In this second embodiment, the same reference numerals are used for components that are the same as in the first embodiment, and their descriptions are omitted.

[0126] In Embodiment 2, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0127] The work terminal 1A includes a communication unit 11A, a control unit 12, a memory 13, an input unit 14, a camera 15, a microphone 16, a speaker 17, and a GPS (Global Positioning System) receiver 18. The following description of the work terminal 1A will explain the differences from the work terminal 1 of Embodiment 1.

[0128] The GPS receiver 18 obtains the current location of the work terminal 1A by receiving GPS signals transmitted from GPS satellites.

[0129] The communication unit 11A transmits location information indicating the current location of the work terminal 1A, acquired by the GPS receiver unit 18, to the server 2.

[0130] In this embodiment 2, the current location of the work terminal 1A is obtained from GPS signals, but the disclosure is not limited thereto, and the current location of the work terminal 1A may be obtained from base station information of a mobile phone or wireless LAN terminal.

[0131] Server 2A comprises a communication unit 21A, a control unit 22A, and a memory 23A. Server 2A is an example of an information processing device. The following description of Server 2A will explain the differences from Server 2 of Embodiment 1.

[0132] The communication unit 21A receives location information from the work terminal 1A indicating the current location of the work terminal 1A. The communication unit 21A receives area designation information from the support terminal 3A used by the supporter, indicating a predetermined area at the work site designated by the supporter located remotely. The communication unit 21A stores the received area designation information in the memory 23A.

[0133] The control unit 22A starts recording the video and first audio to the memory 23A when the location of the work terminal 1A, indicated by the location information, enters a predetermined area indicated by the area designation information. The control unit 22A also stops recording the video and first audio to the memory 23A when the location of the work terminal 1A, indicated by the location information, leaves the predetermined area indicated by the area designation information.

[0134] Furthermore, the control unit 22A may record not only the video and first audio from the work terminal 1A to the memory 23A, but also the video and first audio from the work terminal 1A and the second audio from the support terminal 3A to the memory 23A. That is, the control unit 22A may start recording the video, first audio, and second audio received by the communication unit 21A to the memory 23A when the supporter initiates a recording start operation. Also, the control unit 22A may stop recording the video, first audio, and second audio received by the communication unit 21A to the memory 23A when the supporter initiates a recording end operation.

[0135] Memory 23A stores the area designation information received by the communication unit 21A.

[0136] The support terminal 3A comprises a communication unit 31A, a control unit 32, a memory 33, a display unit 34A, a speaker 35, a microphone 36, and an input unit 37A. The following description of the support terminal 3A will explain the differences from the support terminal 3 of Embodiment 1.

[0137] The communication unit 31A receives drawing information at the work site. The drawing information indicates the location of equipment and other items placed at the work site. The communication unit 31A may receive the drawing information from other terminals or from server 2A.

[0138] The display unit 34A displays the drawing information received by the communication unit 31A.

[0139] The input unit 37A accepts the designation of a predetermined area by the supporter for the drawing information displayed on the display unit 34A. When the work terminal 1A enters the predetermined area, recording of video and audio from the work terminal 1A to the server 2A begins. When the work terminal 1A leaves the predetermined area, recording of video and audio from the work terminal 1A to the server 2A ends. The supporter designates a predetermined area on the drawing of the work site. For example, the input unit 37A accepts the designation of a predetermined area by the supporter enclosing the predetermined area on the drawing of the work site with a line.

[0140] The communication unit 31A transmits area designation information, which indicates a predetermined area at the work site designated by the supporter via the input unit 37A, to the server 2A.

[0141] Next, the work support processing performed by the work terminal 1A, server 2A, and support terminal 3A in Embodiment 2 of this disclosure will be described.

[0142] Figure 7 is a flowchart illustrating the work support processing performed by the work terminal 1A in Embodiment 2 of this disclosure.

[0143] The processes in steps S31 to S33 are the same as those in steps S1 to S3 shown in Figure 2, so their explanation will be omitted.

[0144] Next, in step S34, the GPS receiver 18 obtains the current position of the work terminal 1A by receiving GPS signals transmitted from GPS satellites.

[0145] Next, in step S35, the communication unit 11A transmits location information indicating the current location of the work terminal 1A, which has been acquired by the GPS receiver unit 18, to the server 2.

[0146] The processes in steps S36 to S38 are the same as those in steps S4 to S6 shown in Figure 2, so their explanation will be omitted.

[0147] Figure 8 is a flowchart illustrating the work support processing performed by server 2A in Embodiment 2 of this disclosure.

[0148] First, in step S41, the communication unit 21A receives area designation information indicating a predetermined area at the work site, which is transmitted by the support terminal 3A.

[0149] Next, in step S42, the communication unit 21A stores the received area designation information in the memory 23A.

[0150] The processes in steps S43 to S46 are the same as those in steps S11 to S14 shown in Figure 3, so their explanation will be omitted.

[0151] Next, in step S47, the communication unit 21A receives location information indicating the current location of the work terminal 1A transmitted by the work terminal 1A.

[0152] Next, in step S48, the control unit 22A determines whether the location of the work terminal 1A, indicated by the location information received by the communication unit 21A, is within a predetermined area indicated by the area designation information stored in the memory 23A.

[0153] If it is determined that the position of the work terminal 1A is within a predetermined area (YES in step S48), then in step S49, the control unit 22A determines whether or not the video, first audio, and second audio are being recorded.

[0154] If it is determined that the video, first audio, and second audio are not currently being recorded (NO in step S49), then in step S50, the control unit 22A starts recording the video, first audio, and second audio received by the communication unit 21A into the memory 23A. After that, the process returns to step S43.

[0155] On the other hand, if it is determined that the video, first audio, and second audio are being recorded (YES in step S49), the process returns to step S43.

[0156] Furthermore, if it is determined that the position of the work terminal 1A is not within the predetermined area, that is, if it is determined that the position of the work terminal 1A is outside the predetermined area (NO in step S48), in step S51, the control unit 22A determines whether or not the video, first audio, and second audio are being recorded.

[0157] If it is determined that the video, first audio, and second audio are being recorded (YES in step S51), then in step S52, the control unit 22A terminates recording the video, first audio, and second audio received by the communication unit 21A to the memory 23A. After that, the process returns to step S43.

[0158] On the other hand, if it is determined that the video, first audio, and second audio are not being recorded (NO in step S51), the process returns to step S43.

[0159] Figure 9 is a flowchart illustrating the work support processing by the support terminal 3A in Embodiment 2 of this disclosure.

[0160] First, in step S61, the input unit 37A receives a designation from the supporter of a predetermined area at the work site. The supporter designates the range of a predetermined area on the work site diagram displayed on the display unit 34A where recording to the server 2A will begin when the work terminal 1A enters.

[0161] Next, in step S62, the communication unit 31A transmits area designation information indicating a predetermined area at the work site specified by the input unit 37A to the server 2A.

[0162] The processes in steps S63 to S67 are the same as those in steps S21 to S25 shown in Figure 4, so their explanation will be omitted.

[0163] Thus, when the work terminal 1A enters a predetermined area at the work site designated by a remote supporter, recording of the video and first audio to the memory 23A begins. Therefore, the worker does not need to perform any operation to start recording; the video and first audio can be recorded to the memory 23A simply by the remote supporter designating a predetermined area at the work site.

[0164] In this embodiment 2, the server 2A receives area designation information and determines whether the position of the work terminal 1A is within a predetermined area, but this disclosure is not limited thereto. The work terminal 1A may receive area designation information and determine whether the position of the work terminal 1A is within a predetermined area. If it is determined that the position of the work terminal 1A is within a predetermined area, the work terminal 1A may send a recording start signal to the server 2A to instruct the start of recording. When the server 2A receives the recording start signal sent by the work terminal 1A, it may start recording the video, first audio, and second audio to the memory 23A. Furthermore, if it is determined that the position of the work terminal 1A has left the predetermined area, the work terminal 1A may send a recording end signal to the server 2A to instruct the end of recording. When the server 2A receives the recording end signal sent by the work terminal 1A, it may end the recording of the video, first audio, and second audio to the memory 23A.

[0165] (Embodiment 3) In Embodiment 2, recording of video and first audio to memory 23A is initiated when the work terminal 1A enters a predetermined area at the work site designated by a remote supporter. In Embodiment 3, recording of video and first audio to memory is initiated when the work terminal approaches the work object at the work site designated by a remote supporter.

[0166] Figure 10 shows the configuration of the work support system 10B according to this embodiment 3.

[0167] The work support system 10B shown in Figure 10 comprises a work terminal 1B, a server 2B, and a support terminal 3B. In this embodiment 3, the same reference numerals are used for components that are the same as in embodiment 1, and their descriptions are omitted.

[0168] In Embodiment 3, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0169] Equipment and other work objects placed at the work site are equipped with beacon transmitters that emit beacon signals. A beacon transmitter consists of a transmitter that transmits a beacon signal compliant with the BLE (Bluetooth® Low Energy) communication protocol, for example. The beacon transmitter is installed on the work object and wirelessly transmits a beacon signal containing a beacon ID (identification information) for identifying the work object. The beacon ID is, for example, a UUID (Universally Unique Identifier), MajorID, or MinorID, and uniquely identifies the beacon transmitter as well as the work object equipped with the beacon transmitter. The beacon transmitter transmits a beacon signal at a fixed interval. The beacon transmitter has a memory that stores the beacon ID for identifying the work object in advance and transmits the beacon signal containing the beacon ID. A beacon signal is an example of a wireless signal.

[0170] The work terminal 1B includes a communication unit 11B, a control unit 12, a memory 13, an input unit 14, a camera 15, a microphone 16, a speaker 17, and a beacon receiver 19. The following description of the work terminal 1B will explain the differences from the work terminal 1 of Embodiment 1.

[0171] The beacon receiver 19 receives beacon signals transmitted from a beacon transmitter at the work site. The beacon receiver 19 includes an antenna compliant with the BLE standard. The beacon receiver 19 receives beacon signals transmitted by the beacon transmitter. The beacon receiver 19 also measures the radio wave strength (RSSI (Received Signal Strength Indicator)) of the received beacon signal.

[0172] The communication unit 11B transmits signal information to the server 2B, including the beacon ID contained in the beacon signal received by the beacon receiver unit 19 and the radio wave intensity of the beacon signal measured by the beacon receiver unit 19.

[0173] Furthermore, when the beacon receiver 19 receives multiple beacon signals, it measures the radio wave strength of the multiple beacon signals. The communication unit 11B transmits signal information to the server 2B, which includes multiple beacon IDs contained in each of the multiple beacon signals received by the beacon receiver 19, and the multiple radio wave strengths of each of the multiple beacon signals measured by the beacon receiver 19.

[0174] Server 2B comprises a communication unit 21B, a control unit 22B, and a memory 23B. Server 2B is an example of an information processing device. The following description of Server 2B will explain the differences from Server 2 of Embodiment 1.

[0175] The communication unit 21B receives the beacon ID (identification information) contained in the beacon signal (wireless signal) transmitted from the work target at the work site designated by the supporter in a remote location, from the support terminal 3B used by the supporter. The communication unit 21B stores the beacon ID received from the support terminal 3B in the memory 23B.

[0176] Furthermore, the communication unit 21B receives signal information from the work terminal 1B, including the beacon ID (identification information) contained in the beacon signal (radio signal) received by the work terminal 1B, and the radio wave strength of the beacon signal (radio signal) measured by the work terminal 1B.

[0177] The control unit 22B starts recording the video and first audio to memory 23B when the radio wave strength of the beacon signal (wireless signal) received from work terminal 1B, which includes the beacon ID (identification information) of the work target received from support terminal 3B, is above a threshold. The control unit 22B also stops recording the video and first audio to memory 23B when the radio wave strength of the beacon signal (wireless signal) received from work terminal 1B, which includes the beacon ID (identification information) of the work target received from support terminal 3B, falls below a threshold.

[0178] Furthermore, the control unit 22B may record not only the video and first audio from the work terminal 1B to the memory 23B, but also the video and first audio from the work terminal 1B and the second audio from the support terminal 3B. That is, the control unit 22B may start recording the video, first audio, and second audio received by the communication unit 21B to the memory 23B when the supporter initiates a recording start operation. Also, the control unit 22B may terminate the recording of the video, first audio, and second audio received by the communication unit 21B to the memory 23B when the supporter initiates a recording end operation.

[0179] Memory 23B stores the beacon ID (identification information) of the target of the work received by the communication unit 21B.

[0180] The support terminal 3B comprises a communication unit 31B, a control unit 32, a memory 33, a display unit 34B, a speaker 35, a microphone 36, and an input unit 37B. The following description of the support terminal 3B will explain the differences from the support terminal 3 of Embodiment 1.

[0181] The display unit 34B displays at least one work target at the work site. A beacon ID (identification information) is pre-associated with each of the at least one work targets.

[0182] The input unit 37B accepts a designation by the supporter of a work target from among the at least one work target displayed on the display unit 34B. There is at least one work target at the work site. The supporter designates a work target to be supported from among the at least one work target. When the work terminal 1B approaches the work target designated by the supporter, recording of video and first audio from the work terminal 1B to the server 2B begins. When the work terminal 1B moves away from the work target designated by the supporter, recording of video and first audio from the work terminal 1B to the server 2B ends.

[0183] The communication unit 31B transmits the beacon ID (identification information) associated with the work target specified by the supporter via the input unit 37B to the server 2A.

[0184] Next, the work support processing performed by the work terminal 1B, server 2B, and support terminal 3B in Embodiment 3 of this disclosure will be described.

[0185] Figure 11 is a flowchart illustrating the work support processing by the work terminal 1B in Embodiment 3 of this disclosure.

[0186] The processes in steps S71 to S73 are the same as those in steps S1 to S3 shown in Figure 2, so their explanation will be omitted.

[0187] Next, in step S74, the beacon receiver 19 receives a beacon signal transmitted from a beacon transmitter at the work site.

[0188] Next, in step S75, the beacon receiver 19 measures the radio wave intensity of the received beacon signal.

[0189] Next, in step S76, the communication unit 11B transmits signal information to the server 2B, including the beacon ID contained in the beacon signal received by the beacon receiving unit 19 and the radio wave intensity of the beacon signal measured by the beacon receiving unit 19.

[0190] The processes in steps S77 to S79 ​​are the same as those in steps S4 to S6 shown in Figure 2, so their explanation will be omitted.

[0191] Figure 12 is a flowchart illustrating the work support processing performed by Server 2B in Embodiment 3 of this disclosure.

[0192] First, in step S81, the communication unit 21B receives the beacon ID of the work target at the work site transmitted by the support terminal 3B.

[0193] Next, in step S82, the communication unit 21B stores the received beacon ID of the work target in the memory 23B.

[0194] The processes in steps S83 to S86 are the same as those in steps S11 to S14 shown in Figure 3, so their explanation will be omitted.

[0195] Next, in step S87, the communication unit 21B receives signal information from the work terminal 1B, including the beacon ID included in the beacon signal received by the work terminal 1B and the radio wave intensity of the beacon signal measured by the work terminal 1B.

[0196] Next, in step S88, the control unit 22B determines whether the radio wave intensity of the beacon signal received from the work terminal 1B, including the beacon ID of the work target received from the support terminal 3B, is above a threshold.

[0197] If it is determined that the radio wave intensity of the beacon signal containing the beacon ID of the work target is above a threshold (YES in step S88), then in step S89, the control unit 22B determines whether or not the video, first audio, and second audio are being recorded.

[0198] If it is determined that the video, first audio, and second audio are not currently being recorded (NO in step S89), then in step S90, the control unit 22B starts recording the video, first audio, and second audio received by the communication unit 21B into the memory 23B. After that, the process returns to step S83.

[0199] On the other hand, if it is determined that the video, first audio, and second audio are being recorded (YES in step S89), the process returns to step S83.

[0200] Furthermore, if it is determined that the radio wave strength of the beacon signal including the beacon ID of the target of work is not above a threshold, that is, if it is determined that the radio wave strength of the beacon signal including the beacon ID of the target of work is less than the threshold (NO in step S88), in step S91, the control unit 22B determines whether or not the video, first audio, and second audio are being recorded.

[0201] If it is determined that the video, first audio, and second audio are being recorded (YES in step S91), then in step S92, the control unit 22B terminates recording the video, first audio, and second audio received by the communication unit 21B to the memory 23B. After that, the process returns to step S83.

[0202] On the other hand, if it is determined that the video, first audio, and second audio are not being recorded (NO in step S91), the process returns to step S83.

[0203] Figure 13 is a flowchart illustrating the work support processing by the support terminal 3B in Embodiment 3 of this disclosure.

[0204] First, in step S101, the input unit 37B accepts a designation of a work target by a support worker assisting with the work. The support worker selects a work target from among the at least one work target displayed on the display unit 34B, and when the work terminal 1B approaches, recording to the server 2B will begin.

[0205] Next, in step S102, the communication unit 31B transmits the beacon ID of the work target specified by the input unit 37B to the server 2B.

[0206] The processes in steps S103 to S107 are the same as those in steps S21 to S25 shown in Figure 4, so their explanation will be omitted.

[0207] Thus, when the work terminal 1B approaches the work object at the work site designated by the supporter in a remote location, recording of the video and first audio to the memory 23B begins. Therefore, the worker does not need to perform any operation to start recording; the supporter in a remote location can record the video and first audio to the memory 23B by specifying the identification information of the work object at the work site.

[0208] In this embodiment 3, the server 2B receives the beacon ID of the target of work and determines whether the radio wave strength of the beacon signal including the beacon ID of the target of work is above a threshold, but this disclosure is not limited thereto. The work terminal 1B may receive the beacon ID of the target of work and determine whether the radio wave strength of the beacon signal including the beacon ID of the target of work is above a threshold. If it is determined that the radio wave strength of the beacon signal including the beacon ID of the target of work is above a threshold, the work terminal 1B may send a recording start signal to the server 2B to instruct the start of recording. When the server 2B receives the recording start signal sent by the work terminal 1B, it may start recording the video, first audio, and second audio into memory 23B. If it is determined that the radio wave strength of the beacon signal including the beacon ID of the target of work is below a threshold, the work terminal 1B may send a recording end signal to the server 2B to instruct the end of recording. When the server 2B receives the recording end signal sent by the work terminal 1B, it may end the recording of the video, first audio, and second audio into memory 23B.

[0209] (Embodiment 4) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received, whereas in Embodiment 4, recording of the video and first audio to memory is initiated when a predetermined keyword stored in advance is included in the second audio.

[0210] Figure 14 shows the configuration of the work support system 10C according to this embodiment 4.

[0211] The work support system 10C shown in Figure 14 comprises a work terminal 1, a server 2C, and a support terminal 3C. In this embodiment 4, the same reference numerals are used for components that are the same as in embodiment 1, and their descriptions are omitted.

[0212] In Embodiment 4, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0213] Server 2C comprises a communication unit 21C, a control unit 22C, and a memory 23C. Server 2C is an example of an information processing device. The following description of Server 2C will explain the differences from Server 2 of Embodiment 1.

[0214] The communication unit 21C receives second audio signals from the support terminal 3C, which is used by a support worker in a remote location.

[0215] Memory 23C stores a predetermined start keyword and a predetermined end keyword in advance. The predetermined start keyword is, for example, a demonstrative pronoun such as "there" or "over there," the name of the work object, or the name of the work object's component. The predetermined end keyword is, for example, a phrase to end recording, such as "end recording." The predetermined start keyword and predetermined end keyword may be entered by a supporter. Memory 23C may store one start keyword or multiple start keywords. Similarly, memory 23C may store one end keyword or multiple end keywords.

[0216] The control unit 22C starts recording the video and first audio to memory 23C when a predetermined start keyword, which is pre-stored in memory 23C, is included in the second audio. The control unit 22C also stops recording the video and first audio to memory 23C when a predetermined end keyword, which is pre-stored in memory 23C, is included in the second audio.

[0217] The control unit 22C performs speech recognition on the second audio received by the communication unit 21 and converts the second audio into text. The control unit 22C then determines whether a predetermined start keyword, which is pre-stored in the memory 23C, is included in the converted second audio. If it is determined that the predetermined start keyword is included in the second audio, the control unit 22C starts recording the video and the first audio into the memory 23C.

[0218] Furthermore, if it is determined that a predetermined start keyword is not included in the second audio, the control unit 22C determines whether a predetermined end keyword, which is pre-stored in memory 23C, is included in the transcribed second audio. If it is determined that the predetermined end keyword is included in the second audio, the control unit 22C terminates recording the video and the first audio to memory 23C.

[0219] Furthermore, the control unit 22C may record not only the video and first audio from the work terminal 1 to the memory 23C, but also the video and first audio from the work terminal 1 and the second audio from the support terminal 3C to the memory 23C. That is, the control unit 22C may start recording the video, first audio, and second audio received by the communication unit 21C to the memory 23C when the supporter performs a recording start operation. Also, the control unit 22C may stop recording the video, first audio, and second audio received by the communication unit 21C to the memory 23C when the supporter performs a recording end operation.

[0220] The support terminal 3C comprises a communication unit 31C, a control unit 32C, a memory 33, a display unit 34, a speaker 35, a microphone 36, and an input unit 37C. The following description of the support terminal 3A will explain the differences from the support terminal 3 of Embodiment 1.

[0221] The communication unit 31C receives video footage and first audio recordings from the server 2C at the work site. The communication unit 31C also transmits second audio recordings from the surroundings of the support terminal 3C, collected by the microphone 36, to the server 2C.

[0222] Unlike Embodiment 1, the communication unit 31C does not transmit a recording start signal and a recording end signal to the server 2C. Unlike Embodiment 1, the control unit 32C does not determine whether the recording start button has been pressed. Also, unlike Embodiment 1, the control unit 32C does not determine whether the recording end button has been pressed. Unlike Embodiment 1, the input unit 37C does not include a recording start button and a recording end button.

[0223] Next, the work support processing performed by server 2C and support terminal 3C in Embodiment 4 of this disclosure will be described.

[0224] Figure 15 is a flowchart illustrating the work support processing by Server 2C in Embodiment 4 of this disclosure.

[0225] The processes in steps S121 to S124 are the same as those in steps S11 to S14 shown in Figure 3, so their explanation will be omitted.

[0226] Next, in step S125, the control unit 22C performs speech recognition on the second audio received by the communication unit 21 and converts the second audio into text.

[0227] Next, in step S126, the control unit 22C determines whether or not a predetermined start keyword, which is pre-stored in the memory 23C, is included in the second audio that has been converted into text.

[0228] If it is determined that a predetermined start keyword is included in the second audio (YES in step S126), then in step S127, the control unit 22C starts recording the video, the first audio, and the second audio to the memory 23C. After that, the process returns to step S121.

[0229] On the other hand, if it is determined that a predetermined start keyword is not included in the second audio (NO in step S126), in step S128, the control unit 22C determines whether or not a predetermined end keyword, which is pre-stored in the memory 23C, is included in the transcribed second audio.

[0230] If it is determined that the predetermined termination keyword is not included in the second audio (NO in step S128), the process returns to step S121.

[0231] On the other hand, if it is determined that a predetermined termination keyword is included in the second audio (YES in step S128), in step S129, the control unit 22C determines whether or not the video, the first audio, and the second audio are being recorded.

[0232] If it is determined that the video, first audio, and second audio are not being recorded (NO in step S129), the process returns to step S121.

[0233] On the other hand, if it is determined that the video, first audio, and second audio are being recorded (YES in step S129), in step S130, the control unit 22C terminates recording the video, first audio, and second audio to the memory 23C. Then, the process returns to step S121.

[0234] Note that the work support processing by the support terminal 3C in this embodiment 4 is the same as the processing in steps S21 to S25 shown in Figure 4, so the explanation will be omitted.

[0235] In this way, when a supporter located remotely utters a predetermined keyword that has been stored in advance, the recording of the video and first audio to memory 23C begins. Therefore, the operator does not need to perform any operation to start recording; the video and first audio can be recorded to memory 23C simply by the supporter located remotely uttering the predetermined keyword.

[0236] In this embodiment 4, the control unit 22C determines whether a predetermined termination keyword is included in the second audio, but this disclosure is not limited thereto. The support terminal 3C may accept the press of the recording termination button by the supporter. If the recording termination button is pressed by the supporter, the support terminal 3C may send a recording termination signal to the server 2C. The communication unit 21C of the server 2C may receive the recording termination signal sent by the support terminal 3C. The control unit 22C may determine whether the recording termination signal sent by the support terminal 3C has been received. If it is determined that the recording termination signal has been received, the control unit 22C may terminate the recording of the video, the first audio, and the second audio to the memory 23C.

[0237] Furthermore, in this embodiment 4, the server 2C determines whether a predetermined start keyword is included in the second audio, but this disclosure is not limited thereto. The support terminal 3C may determine whether a predetermined start keyword is included in the second audio. If it is determined that the predetermined start keyword is included in the second audio, the support terminal 3C may send a recording start signal to the server 2C to instruct the start of recording. When the server 2C receives the recording start signal sent by the support terminal 3C, it may start recording the video, first audio, and second audio to the memory 23C. The support terminal 3C may also determine whether a predetermined end keyword is included in the second audio. If it is determined that the predetermined end keyword is included in the second audio, the support terminal 3C may send a recording end signal to the server 2C to instruct the end of recording. When the server 2C receives the recording end signal sent by the support terminal 3C, it may end recording the video, first audio, and second audio to the memory 23C.

[0238] (Embodiment 5) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 5, however, a speech segment spoken by the supporter in the second audio is detected, and recording of the video and first audio to memory is initiated when the speech segment is detected.

[0239] Figure 16 shows the configuration of the work support system 10D according to this embodiment 5.

[0240] The work support system 10D shown in Figure 16 comprises a work terminal 1, a server 2D, and a support terminal 3C. In this embodiment 5, the same reference numerals are used for components that are the same as those in embodiments 1 and 4, and their descriptions are omitted.

[0241] In Embodiment 5, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0242] Server 2D comprises a communication unit 21D, a control unit 22D, and a memory 23. Server 2D is an example of an information processing device. The following description of Server 2D will explain the differences from Server 2 of Embodiment 1.

[0243] The communication unit 21D receives second audio signals from the support terminal 3C, which is used by a support worker in a remote location.

[0244] The control unit 22D detects the speech intervals in the second audio where the supporter has spoken. The control unit 22D detects speech intervals using general voice activity detection (VAD) techniques. For example, the control unit 22D detects whether a frame is a speech interval based on its amplitude and zero crossing count in a frame composed of the time series of the input second audio. Alternatively, for example, the control unit 22D may calculate the probability that the supporter is speaking using a speech model and the probability that the supporter is not speaking using a noise model, and determine that an interval where the probability obtained from the speech model is higher than the probability obtained from the noise model is a speech interval.

[0245] The control unit 22D starts recording the video and first audio to the memory 23 when a speech interval is detected. The control unit 22D also stops recording the video and first audio to the memory 23 when a speech interval is no longer detected.

[0246] The control unit 22D determines whether the second audio is in a speech interval. If it is determined that the second audio is in a speech interval, the control unit 22D starts recording the video and the first audio to the memory 23. If, after it has been determined that the second audio is in a speech interval, it is determined that the second audio is not in a speech interval, the control unit 22D stops recording the video and the first audio to the memory 23.

[0247] Furthermore, the control unit 22D may record not only the video and first audio from the work terminal 1 to the memory 23, but also the video and first audio from the work terminal 1 and the second audio from the support terminal 3C to the memory 23. That is, the control unit 22C may start recording the video, first audio, and second audio received by the communication unit 21D to the memory 23 when the supporter initiates a recording start operation. Also, the control unit 22D may stop recording the video, first audio, and second audio received by the communication unit 21D to the memory 23 when the supporter initiates a recording end operation.

[0248] Next, the work support processing by Server 2C in Embodiment 5 of this disclosure will be described.

[0249] Figure 17 is a flowchart illustrating the work support processing by Server 2C in Embodiment 5 of this disclosure.

[0250] The processes in steps S141 to S144 are the same as those in steps S11 to S14 shown in Figure 3, so their explanation will be omitted.

[0251] Next, in step S145, the control unit 22D detects the speech interval in the second voice.

[0252] Next, in step S146, the control unit 22D determines whether the second voice is in a speech interval.

[0253] If it is determined that the second audio is in a speech interval (YES in step S146), then in step S147, the control unit 22D determines whether the video, the first audio, and the second audio are being recorded.

[0254] If it is determined that the video, first audio, and second audio are being recorded (YES in step S147), the process returns to step S141.

[0255] On the other hand, if it is determined that the video, first audio, and second audio are not currently being recorded (NO in step S147), in step S148, the control unit 22D starts recording the video, first audio, and second audio to the memory 23. Then, the process returns to step S141.

[0256] On the other hand, if it is determined that the second audio is not in a speech interval (NO in step S146), in step S149, the control unit 22D determines whether the video, the first audio, and the second audio are being recorded.

[0257] If it is determined that the video, first audio, and second audio are not being recorded (NO in step S149), the process returns to step S141.

[0258] On the other hand, if it is determined that the video, first audio, and second audio are being recorded (YES in step S149), in step S150, the control unit 22D terminates recording the video, first audio, and second audio to the memory 23. Then, the process returns to step S141.

[0259] In this way, when a supporter in a remote location speaks, recording of the video and the first audio to memory 23 begins. Therefore, the worker does not need to perform any operation to start recording; the video and the first audio can be recorded to memory 23 simply by the supporter in a remote location speaking.

[0260] In this embodiment 5, the control unit 22D detects the utterance section spoken by the supporter in the second audio, but this disclosure is not limited thereto. The control unit 22D may also detect the utterance sections spoken by the worker and supporter in the first and second audio. The control unit 22D may determine whether the first and second audio are utterance sections. If it is determined that the first and second audio are utterance sections, the control unit 22D may start recording the video and the first audio to the memory 23. Furthermore, if it is determined that the first and second audio are not utterance sections after it has been determined that they are utterance sections, the control unit 22D may stop recording the video and the first audio to the memory 23.

[0261] Furthermore, in this embodiment 5, the server 2D detects the utterance section spoken by the supporter in the second audio and determines whether or not the second audio is an utterance section, but this disclosure is not particularly limited thereto. The support terminal 3C may also detect the utterance section spoken by the supporter in the second audio and determine whether or not the second audio is an utterance section. If it is determined that the second audio is an utterance section, the support terminal 3C may send a recording start signal to the server 2D to instruct the start of recording. When the server 2D receives the recording start signal sent by the support terminal 3C, it may start recording the video, the first audio, and the second audio to the memory 23. Also, if it is determined that the second audio is not an utterance section after it has been determined that the second audio is an utterance section, the support terminal 3C may send a recording end signal to the server 2D to instruct the end of recording. When the server 2D receives the recording end signal sent by the support terminal 3C, it may end recording the video, the first audio, and the second audio to the memory 23.

[0262] Furthermore, in this embodiment 5, the control unit 22D may detect the utterance section spoken by the supporter in the second audio, and trigger the recording of the video and the first audio to the memory 23 when a predetermined keyword stored in advance is included in the detected utterance section of the second audio. In this case, if it is determined in step S146 of Figure 17 that the second audio is an utterance section, the processing in steps S125 to S130 of Figure 15 may be performed.

[0263] That is, the control unit 22D determines whether the second voice is in the speaking section. When it is determined that the second voice is in the speaking section, the control unit 22D may perform voice recognition on the second voice received by the communication unit 21D and convert the second voice into text. Then, the control unit 22D may determine whether a predetermined start keyword stored in advance in the memory 23 is included in the text-converted second voice. When it is determined that the predetermined start keyword is included in the second voice, the control unit 22D may start recording the moving image and the first voice in the memory 23. Also, during the recording of the moving image and the first voice, when it is determined that the second voice is not in the speaking section, the control unit 22D may end the recording of the moving image and the first voice in the memory 23.

[0264] (Embodiment 6) In Embodiment 1, the recording of the moving image and the first voice in the memory 23 is started triggered by the reception of the recording start signal. However, in Embodiment 6, from the received moving image, the actions of the supporter at the work site are recognized, and the recording of the moving image and the first voice in the memory is started triggered by the fact that the recognized actions are predetermined actions.

[0265] FIG. 18 is a diagram showing the configuration of the work support system 10E according to the sixth embodiment.

[0266] The work support system 10E shown in FIG. 18 includes a work terminal 1E and a server 2E. In the sixth embodiment, the same components as those in Embodiment 1 are denoted by the same reference numerals, and the description thereof is omitted.

[0267] In the sixth embodiment, the worker performing the work is at the work site, and the supporter who supports the work of the worker is also at the work site. The worker takes a picture of the state in which the supporter is supporting the work on the work target at the work site.

[0268] The work terminal 1E includes a communication unit 11E, a control unit 12E, a memory 13, an input unit 14E, a camera 15, and a microphone 16. In the following description of the work terminal 1E, the differences from the work terminal 1 in Embodiment 1 will be described.

[0269] The communication unit 11E transmits the moving image captured by the camera 15 and the first audio collected by the microphone 16 to the server 2E.

[0270] The input unit 14E includes a recording end button for ending the recording of the moving image and the first audio to the server 2E. The control unit 12E determines whether the recording end button of the input unit 14E has been pressed. When the recording end button is pressed by the operator, the communication unit 11E transmits a recording end signal instructing the end of recording to the server 2E.

[0271] The server 2E includes a communication unit 21E, a control unit 22E, and a memory 23E. The server 2E is an example of an information processing device. In the following description of the server 2E, the differences from the server 2 in Embodiment 1 will be described.

[0272] The communication unit 21E receives the moving image captured at the work site and the first audio collected at the work site from the work terminal 1E used by the worker at the work site.

[0273] The control unit 22E recognizes the actions of the supporter at the work site from the moving image received by the communication unit 21E. The control unit 22E starts recording the moving image and the first audio to the memory 23E using as a trigger that the recognized action of the supporter is a predetermined action. The predetermined action is an action in which the supporter points a finger at the work target.

[0274] The control unit 22E estimates the skeleton of the person shown in the moving image using a learned neural network. Also, the control unit 22E recognizes the action of the person from the estimated skeleton using a learned neural network.

[0275] The control unit 22E determines whether the recognized action of the support worker is a predetermined action. If it is determined that the recognized action of the support worker is a predetermined action, the control unit 22E starts recording the video and first audio to the memory 23E. For example, when a support worker assists with work at a work site, they point their finger towards the work object. The control unit 22E recognizes the support worker's pointing action, determines that the support worker has started assisting with the work, and starts recording the video and first audio to the memory 23E. The first audio includes the worker's voice and the support worker's voice.

[0276] Furthermore, the communication unit 21E receives a recording termination signal from the work terminal 1E, which instructs the end of recording based on input operations by the worker at the work site. The control unit 22E terminates the recording of the video and first audio to the memory 23E, triggered by the communication unit 21E receiving the recording termination signal.

[0277] Next, the work support processing performed by the work terminal 1E and the server 2E in Embodiment 6 of this disclosure will be described.

[0278] Figure 19 is a flowchart illustrating the work support processing performed by the work terminal 1E in Embodiment 6 of this disclosure.

[0279] The processes in steps S151 to S153 are the same as those in steps S1 to S3 shown in Figure 2, so their explanation will be omitted.

[0280] Next, in step S154, the control unit 12E determines whether or not the recording end button on the input unit 14E has been pressed.

[0281] If it is determined that the recording stop button has not been pressed (NO in step S154), the process returns to step S151.

[0282] On the other hand, when it is determined that the recording end button has been pressed (YES in step S154), in step S155, the communication unit 11E transmits a recording end signal instructing the end of recording to the server 2E.

[0283] The process of step S156 is the same as the process of step S6 shown in FIG. 2, so the description thereof is omitted.

[0284] FIG. 20 is a flowchart for explaining the work support process by the server 2E in Embodiment 6 of the present disclosure.

[0285] First, in step S161, the communication unit 21E receives the moving image and the first audio transmitted by the work terminal 1E.

[0286] Next, in step S162, the control unit 22E recognizes the operation of the supporter at the work site from the moving image received by the communication unit 21E.

[0287] Next, in step S163, the control unit 22E determines whether the recognized operation is a predetermined operation determined in advance.

[0288] Here, when it is determined that the recognized operation is a predetermined operation (YES in step S163), in step S164, the control unit 22E determines whether the moving image and the first audio are being recorded.

[0289] Here, when it is determined that the moving image and the first audio are being recorded (YES in step S164), the process returns to step S161.

[0290] On the other hand, when it is determined that the moving image and the first audio are not being recorded (NO in step S164), in step S165, the control unit 22E starts recording the moving image and the first audio received by the communication unit 21E in the memory 23E. Then, the process returns to step S161.

[0291] On the other hand, if it is determined that the recognized operation is not a predetermined operation (NO in step S163), in step S166, the control unit 22E determines whether or not a recording termination signal instructing the end of recording has been received by the communication unit 21E based on an input operation by the worker at the work site.

[0292] If it is determined that a recording completion signal has been received (YES in step S166), in step S167, the control unit 22E terminates the recording of the video and first audio received by the communication unit 21E to the memory 23E. After that, the process returns to step S161.

[0293] On the other hand, if it is determined that no recording completion signal has been received (NO in step S166), the process returns to step S161.

[0294] In this way, when a support person at the work site performs a predetermined action, recording of the video and first audio to memory 23E begins. Therefore, the worker does not need to perform any operation to start recording; the video and first audio can be recorded to memory 23E simply by the support person at the work site performing the predetermined action.

[0295] (Embodiment 7) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 7, recording of the video and first audio to memory is initiated when a predetermined keyword, which is stored in advance, is included in the first audio, which includes the voice of a support worker at the work site.

[0296] Figure 21 shows the configuration of the work support system 10F according to this embodiment 7.

[0297] The work support system 10F shown in Figure 21 includes a work terminal 1F and a server 2F. In this embodiment 7, the same reference numerals are used for components that are the same as in embodiment 1, and their descriptions are omitted.

[0298] In Embodiment 7, the worker performing the work is at the work site, and the support worker assisting the worker is also at the work site. At the work site, the worker uses work terminal 1F to film the support worker assisting the work on the object of work.

[0299] The work terminal 1F comprises a communication unit 11F, a control unit 12, a memory 13, an input unit 14, a camera 15, and a microphone 16. The following description of the work terminal 1F will explain the differences from the work terminal 1 of Embodiment 1.

[0300] The communications unit 11F transmits the video footage captured by the camera 15 and the first audio recording collected by the microphone 16 to the server 2F.

[0301] Server 2F comprises a communication unit 21F, a control unit 22F, and a memory 23F. Server 2F is an example of an information processing device. The following description of Server 2F will explain the differences from Server 2 of Embodiment 1.

[0302] The communications unit 21F receives video footage and audio recordings taken at the work site from the work terminal 1F used by the workers at the work site.

[0303] Memory 23F pre-stores a predetermined start keyword and a predetermined end keyword. The predetermined start keyword is, for example, a demonstrative pronoun such as "there" or "over there," the name of the work object, or the name of the work object's component. The predetermined end keyword is, for example, a phrase used to end recording, such as "end recording." The predetermined start keyword and predetermined end keyword may be entered by a support worker. Memory 23F may store one start keyword or multiple start keywords. Similarly, memory 23F may store one end keyword or multiple end keywords.

[0304] The control unit 22F starts recording the video and first audio to memory 23F when it detects that a predetermined keyword pre-stored in memory 23F is included in the first audio, which includes the voice of a support worker at the work site. The control unit 22F also stops recording the video and first audio to memory 23F when it detects that a predetermined termination keyword pre-stored in memory 23F is included in the first audio.

[0305] The control unit 22F performs speech recognition on the first audio received by the communication unit 21F and converts the first audio into text. The control unit 22F then determines whether a predetermined start keyword, which is pre-stored in the memory 23F, is included in the converted first audio. If it is determined that the predetermined start keyword is included in the first audio, the control unit 22F starts recording the video and the first audio into the memory 23F.

[0306] For example, when an assistant provides support for work at a work site, they converse with a worker wearing a work terminal 1F. At this time, the assistant speaks a predetermined start keyword when they begin recording video and audio to server 2F. If the predetermined start keyword is included in the first audio collected at the work site, the control unit 22F determines that the assistant has started providing support for the work and begins recording video and audio to memory 23F.

[0307] Furthermore, if it is determined that a predetermined start keyword is not included in the first audio, the control unit 22F determines whether a predetermined end keyword, which is pre-stored in memory 23F, is included in the transcribed first audio. If it is determined that the predetermined end keyword is included in the first audio, the control unit 22F terminates the recording of the video and the first audio to memory 23F. For example, the supporter speaks a predetermined end keyword at the time when they terminate the recording of the video and the first audio to server 2F.

[0308] Next, the work support processing performed by the work terminal 1F and server 2F in Embodiment 7 of this disclosure will be described.

[0309] Note that the work support processing by the work terminal 1F in this embodiment 7 is the same as the processing in steps S1, S2, S3, and S6 shown in Figure 2, so the explanation will be omitted.

[0310] Figure 22 is a flowchart illustrating the work support processing performed by server 2F in Embodiment 7 of this disclosure.

[0311] First, in step S171, the communication unit 21F receives the video and audio transmitted by the work terminal 1F.

[0312] Next, in step S172, the control unit 22F performs speech recognition on the first audio received by the communication unit 21F and converts the first audio into text.

[0313] Next, in step S173, the control unit 22F determines whether or not a predetermined start keyword, which is pre-stored in the memory 23F, is included in the first audio that has been converted into text.

[0314] If it is determined that a predetermined start keyword is included in the first audio (YES in step S173), then in step S174, the control unit 22F starts recording the video and the first audio to the memory 23F. After that, the process returns to step S171.

[0315] On the other hand, if it is determined that a predetermined start keyword is not included in the first audio (NO in step S173), in step S175, the control unit 22F determines whether or not a predetermined end keyword, which is pre-stored in the memory 23F, is included in the first audio that has been transcribed into text.

[0316] If it is determined that the predetermined termination keyword is not included in the first audio (NO in step S175), the process returns to step S171.

[0317] On the other hand, if it is determined that a predetermined termination keyword is included in the first audio (YES in step S175), in step S176, the control unit 22F determines whether or not the video and the first audio are being recorded.

[0318] If it is determined that the video and first audio are not being recorded (NO in step S176), the process returns to step S171.

[0319] On the other hand, if it is determined that the video and first audio are being recorded (YES in step S176), in step S177, the control unit 22F terminates recording the video and first audio to memory 23F. Then, the process returns to step S171.

[0320] In this way, when a support person at the work site utters a predetermined keyword that has been stored in advance, the recording of the video and the first audio to memory 23F begins. Therefore, the worker does not need to perform any operation to start recording; the video and the first audio can be recorded to memory 23F simply by the support person at the work site uttering the predetermined keyword.

[0321] In this embodiment 7, the control unit 22F determines whether a predetermined termination keyword is included in the first audio, but this disclosure is not limited thereto. The work terminal 1F may accept the operator pressing the recording termination button. When the operator presses the recording termination button, the work terminal 1F may send a recording termination signal to the server 2F. The communication unit 21F of the server 2F may receive the recording termination signal sent by the work terminal 1F. The control unit 22F may determine whether the recording termination signal sent by the work terminal 1F has been received. If it is determined that the recording termination signal has been received, the control unit 22F may terminate the recording of the video and the first audio to the memory 23F.

[0322] Furthermore, in this embodiment 7, the server 2F determines whether or not a predetermined start keyword is included in the first audio, but this disclosure is not particularly limited thereto. The work terminal 1F may determine whether or not a predetermined start keyword is included in the first audio. If it is determined that the predetermined start keyword is included in the first audio, the work terminal 1F may send a recording start signal to the server 2F to instruct the start of recording. When the server 2F receives the recording start signal sent by the work terminal 1F, it may start recording the video and the first audio to the memory 23F. The work terminal 1F may also determine whether or not a predetermined end keyword is included in the first audio. If it is determined that the predetermined end keyword is included in the first audio, the work terminal 1F may send a recording end signal to the server 2F to instruct the end of recording. When the server 2F receives the recording end signal sent by the work terminal 1F, it may end recording the video and the first audio to the memory 23F.

[0323] (Embodiment 8) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 8, recording of the video and first audio to memory is initiated when a still image extracted from the video by a remote supporter is received from a support terminal.

[0324] Figure 23 is a diagram showing the configuration of the work support system 10G according to this embodiment 8.

[0325] The work support system 10G shown in Figure 23 comprises a work terminal 1G, a server 2G, and a support terminal 3G. In this embodiment 8, the same reference numerals are used for components that are the same as in embodiment 1, and their descriptions are omitted.

[0326] In Embodiment 8, the worker performing the task is at the work site, while the support worker assisting the worker is in a remote location.

[0327] The support terminal 3G comprises a communication unit 31G, a control unit 32G, a memory 33, a display unit 34G, a speaker 35, a microphone 36, and an input unit 37G. The following description of the support terminal 3G will explain the differences from the support terminal 3 of Embodiment 1.

[0328] The input unit 37G includes a capture start button for extracting a still image from a moving image displayed by the display unit 34G. The capture start button may be a button that is physically pressed by an assistant, or it may be a button that is displayed on the display unit 34G and clicked with a mouse.

[0329] The control unit 32G determines whether or not the capture start button has been pressed. When the capture start button is pressed by the supporter, the control unit 32G extracts a still image from the video. The display unit 34G displays the still image extracted from the video, and the communication unit 31G periodically transmits the still image extracted from the video to the work terminal 1G via the server 2G. For example, the supporter presses the capture start button when they see a part that requires assistance while watching the video. As a result, the still image at the time the capture start button was pressed is displayed on the display unit 34G and periodically transmitted by the communication unit 31G to the work terminal 1G via the server 2G.

[0330] Furthermore, the input unit 37G receives instruction information such as characters and symbols from the supporter regarding the still image displayed on the display unit 34G. For example, the supporter may draw arrows or write characters on the displayed still image to give specific instructions for a task. The communication unit 31G periodically transmits the still image, overlaid with the instruction information entered by the supporter, to the work terminal 1G via the server 2G.

[0331] Furthermore, the input unit 37G includes a capture termination button for ending the display and transmission of the extracted still image. The capture termination button may be a button that is physically pressed by an assistant, or it may be a button that is displayed on the display unit 34G and clicked with a mouse.

[0332] The control unit 32G determines whether or not the capture end button has been pressed. When the capture end button is pressed by the supporter, the display unit 34G ends the display of the still image, and the communication unit 31G ends the transmission of the still image to the server 2G.

[0333] Server 2G comprises a communication unit 21G, a control unit 22G, and memory 23G. Server 2G is an example of an information processing device. The following description of Server 2G will explain the differences from Server 2 of Embodiment 1.

[0334] The communication unit 21G receives still images extracted by the supporter from the moving images displayed by the support terminal 3G's display unit 34G. The communication unit 21G also transmits the still images received from the support terminal 3G to the work terminal 1G.

[0335] The control unit 22G starts recording the video and first audio to memory 23G when it receives a still image. The control unit 22G determines whether or not a still image has been received by the communication unit 21G. If the communication unit 21G determines that a still image has been received, the control unit 22G starts recording the video and first audio to memory 23G.

[0336] Furthermore, after recording of the video and first audio has started, the control unit 22G terminates recording of the video and first audio to memory 23G when it stops receiving still images. If the communication unit 21G determines that no still images have been received during the recording of the video and first audio, the control unit 22G terminates recording of the video and first audio to memory 23G.

[0337] Furthermore, the control unit 22G may record not only the video and first audio from the work terminal 1 to the memory 23G, but also the video and first audio from the work terminal 1 and the second audio from the support terminal 3G to the memory 23G. That is, the control unit 22G may start recording the video, first audio, and second audio received by the communication unit 21G to the memory 23G when the supporter initiates a recording start operation. Also, the control unit 22G may terminate the recording of the video, first audio, and second audio received by the communication unit 21G to the memory 23G when the supporter initiates a recording end operation.

[0338] Furthermore, the control unit 22G may record the video and first audio from the work terminal 1 and the second audio and still images from the support terminal 3G into the memory 23G. That is, the control unit 22G may start recording the video, first audio, second audio, and still images received by the communication unit 21G into the memory 23G as a trigger when it receives a still image. Also, the control unit 22G may stop recording the video, first audio, second audio, and still images received by the communication unit 21G into the memory 23G as a trigger when it stops receiving still images while recording the video, first audio, second audio, and still images.

[0339] Furthermore, the control unit 22G may record only the second audio and still images from the support terminal 3G into memory 23, without recording the video and first audio from the work terminal 1 into memory 23G. That is, the control unit 22G may start recording the second audio and still images received by the communication unit 21G into memory 23G as a trigger when it receives a still image. Also, the control unit 22G may stop recording the second audio and still images received by the communication unit 21G into memory 23G as a trigger when it stops receiving still images while recording the second audio and still images.

[0340] Memory 23G may not only non-temporarily record the video and first audio from the work terminal 1, but may also non-temporarily record the video and first audio from the work terminal 1 and the second audio from the support terminal 3. In other words, memory 23 may non-temporarily record the video, first audio, and second audio received by the communication unit 21.

[0341] Furthermore, memory 23G may non-temporarily record video and audio from the work terminal 1G, and second audio and still images from the support terminal 3G. That is, memory 23 may non-temporarily record video, audio, first audio, second audio, and still images received by the communication unit 21G.

[0342] The work terminal 1G comprises a communication unit 11G, a control unit 12G, a memory 13, an input unit 14, a camera 15, a microphone 16, a speaker 17, and a display unit 20. The following description of the work terminal 1G will explain the differences from the work terminal 1 of Embodiment 1.

[0343] The communication unit 11G periodically receives still images extracted from video footage by supporters from server 2G.

[0344] The display unit 20 displays still images received by the communication unit 11G. This allows the worker to receive work assistance from the supporter while viewing still images extracted from moving images by the supporter. The display unit 20 also displays still images with characters and symbols superimposed by the supporter. This allows the worker to receive more detailed work assistance from the supporter while viewing still images with characters and symbols superimposed. The work terminal 1G may also be equipped with a touch panel that integrates the input unit 14 and the display unit 20.

[0345] The display unit 20 may also display moving images captured by the camera 15.

[0346] Next, the work support processing performed by the work terminal 1G, server 2G, and support terminal 3G in Embodiment 8 of this disclosure will be described.

[0347] Figure 24 is a flowchart illustrating the work support processing by the work terminal 1G in Embodiment 8 of this disclosure.

[0348] The processes in steps S181 to S185 are the same as those in steps S1 to S5 shown in Figure 2, so their explanation will be omitted.

[0349] Next, in step S186, the control unit 12G determines whether or not a still image has been received by the communication unit 11G. The communication unit 11G receives the still image transmitted by the server 2G.

[0350] If it is determined that a still image has been received (YES in step S186), then in step S187, the display unit 20 displays the still image received by the communication unit 11G.

[0351] On the other hand, if it is determined that no still image has been received (NO in step S186), the process proceeds to step S188. Furthermore, if it is determined that no still image has been received while a still image is being displayed, the display unit 20 terminates the display of the still image.

[0352] The process in step S188 is the same as the process in step S6 shown in Figure 2, so its explanation is omitted.

[0353] Figure 25 is a flowchart illustrating the work support processing by server 2G in Embodiment 8 of this disclosure.

[0354] The processes in steps S191 to S194 are the same as those in steps S11 to S14 shown in Figure 3, so their explanation will be omitted.

[0355] Next, in step S195, the control unit 22G determines whether or not a still image has been received by the communication unit 21G. The communication unit 21G receives the still image transmitted by the support terminal 3G.

[0356] If it is determined that a still image has been received (YES in step S195), then in step S196, the communication unit 21G transmits the received still image to the work terminal 1G.

[0357] Next, in step S197, the control unit 22G determines whether or not the moving image, first audio, second audio, and still image are being recorded.

[0358] If it is determined that the video, first audio, second audio, and still images are not being recorded (NO in step S197), then in step S198, the control unit 22G starts recording the video, first audio, second audio, and still images received by the communication unit 21G into the memory 23G. After that, the process returns to step S191.

[0359] On the other hand, if it is determined that video, first audio, second audio, and still images are being recorded (YES in step S197), the process returns to step S191.

[0360] Furthermore, if it is determined that no still image has been received (NO in step S195), in step S199, the control unit 22G determines whether or not a moving image, first audio, second audio, and still image are being recorded.

[0361] If it is determined that the video, first audio, second audio, and still image are being recorded (YES in step S199), then in step S200, the control unit 22G terminates recording the video, first audio, second audio, and still image received by the communication unit 21G to the memory 23G. After that, the process returns to step S191.

[0362] On the other hand, if it is determined that the video, first audio, second audio, and still images are not being recorded (NO in step S199), the process returns to step S191.

[0363] Figure 26 is a flowchart illustrating the work support processing by the support terminal 3G in Embodiment 8 of this disclosure.

[0364] The processes in steps S211 to S215 are the same as those in steps S21 to S25 shown in Figure 4, so their explanation will be omitted.

[0365] Next, in step S216, the control unit 32G determines whether or not the capture start button of the input unit 37G has been pressed.

[0366] If it is determined that the capture start button has been pressed (YES in step S216), then in step S217, the control unit 32G extracts a still image from the video received by the communication unit 31G.

[0367] Next, in step S218, the display unit 34G displays the still image extracted by the control unit 32G.

[0368] Next, in step S219, the input unit 37G receives instruction information such as characters and symbols from the supporter regarding the still image displayed on the display unit 34G.

[0369] Next, in step S220, the communication unit 31G transmits the still image extracted from the video to the server 2G. Then, the process returns to step S211. If instruction information such as characters and symbols is entered by the supporter, the communication unit 31G transmits the still image with the instruction information superimposed to the server 2G. The communication unit 31G also transmits the still image to the server 2G with the work terminal 1G as the destination. As a result, the still image is transmitted to the work terminal 1G via the server 2G.

[0370] On the other hand, if it is determined that the capture start button has not been pressed (NO in step S216), in step S221, the control unit 32G determines whether or not a still image is currently being displayed on the display unit 34G.

[0371] If it is determined that a still image is not currently being displayed (NO in step S221), the process returns to step S211.

[0372] On the other hand, if it is determined that a still image is being displayed (YES in step S221), in step S222, the control unit 32G determines whether or not the capture end button of the input unit 37G has been pressed.

[0373] If it is determined that the capture end button has not been pressed (NO in step S222), the process proceeds to step S218.

[0374] On the other hand, if it is determined that the capture end button has been pressed (YES in step S222), in step S223, the display unit 34G ends the display of the still image.

[0375] Next, in step S224, the communication unit 31G terminates the transmission of the still image to the server 2G.

[0376] In this way, when a support worker in a remote location extracts still images to be used to support the work from the video displayed on the display unit 34G of the support terminal 3G, recording of the video and the first audio to memory 23G begins. Therefore, the worker does not need to perform any operation to start recording, and the video and the first audio can be recorded to memory 23G by the support worker in a remote location extracting still images from the video.

[0377] Figure 27 shows an example of a screen displayed on the display unit 34G of the support terminal 3G in this embodiment 8.

[0378] The display unit 34G displays a video 351 of the work site, a capture start button 352, and a capture end button 353. When the assistant operates the mouse, the pointer displayed on the display unit 34G moves over the capture start button 352. When the assistant clicks the mouse button, a still image 354 is extracted from the video 351 and displayed on the display unit 34G. The extracted still image 354 is then sent to the server 2G. As a result, the server 2G starts recording the video, first audio, second audio, and still image.

[0379] Furthermore, while the still image 354 is being displayed, the supporter moves the pointer displayed on the display unit 34G over the capture termination button 353 by operating the mouse. When the supporter clicks the mouse button, the display of the still image 354 ends, and the transmission of the still image 354 also ends. As a result, the server 2G ends recording the video, first audio, second audio, and still image. After the display of the still image 354 ends, the video 351 is displayed.

[0380] Furthermore, the input unit 37G receives instruction information such as characters 355 and symbols 356 from the supporter regarding the still image 354 displayed on the display unit 34G. The supporter uses a mouse or keyboard to write the characters 355 and symbols 356 onto the still image 354 displayed on the display unit 34G. In Figure 27, the characters 355 "Turn" and the symbol 356 representing an arrow are written. Once the instruction information is input, the communication unit 31G transmits the still image 354 with the instruction information superimposed to the server 2G.

[0381] In the example shown in Figure 27, when the capture start button 352 is pressed, the display unit 34G displays only the still image 354, but this disclosure is not limited to this. The display unit 34G may also display the moving image 351 superimposed on the still image 354. For example, the display unit 34G may display the still image 354 in full screen and the moving image 351 in a small size in the lower right portion of the screen.

[0382] Figure 28 shows an example of a screen displayed on the display unit 20 of the work terminal 1G in this embodiment 8.

[0383] The display unit 20 displays a video 201 of the work site and a still image 202 transmitted by the support terminal 3G. The video 201 is a video captured in real time by the camera 15. The display unit 20 displays the video 201 in full screen and the still image 202 in a small size in the lower right corner of the screen. When instruction information is input for the still image 202, the still image 202 with the instruction information superimposed is displayed. By performing the work while viewing the still image 202 displayed on the display unit 20, the worker can receive support from the support worker.

[0384] The control unit 12G of the work terminal 1G may determine whether or not a still image 202 is included in the video 201 captured by the camera 15. If it is determined that a still image 202 is included in the video 201, the display unit 20 may highlight the area 203 in the video 201 that corresponds to the still image 202. In Figure 28, the area 203 that corresponds to the still image 202 is surrounded by a line of a predetermined color. The predetermined color is, for example, red.

[0385] In Figure 28, the display unit 20 displays both a moving image 201 and a still image 202. However, this disclosure is not limited to this, and the display unit 20 may display only the still image 202.

[0386] Furthermore, in this embodiment 8, the control unit 22G starts recording the video and the first audio to memory 23G triggered by the reception of a still image extracted from the video from the support terminal 3G, but the disclosure is not particularly limited thereto. The control unit 22G may also start recording the video and the first audio to memory triggered by the superimposition of instruction information onto the still image. In this case, the control unit 22G does not start recording simply by receiving a still image, but starts recording only when instruction information is superimposed onto the still image. That is, after receiving a still image, the control unit 22G may determine whether or not instruction information has been superimposed onto the still image. If it is determined that instruction information has been superimposed onto the still image, the control unit 22G may start recording the video and the first audio to memory.

[0387] Specifically, in step S197 of Figure 25, if it is determined that the video, first audio, second audio, and still image are not being recorded (NO in step S197), the control unit 22G determines whether or not instruction information has been superimposed on the still image. If it is determined that the instruction information has not been superimposed on the still image, the process returns to step S191. On the other hand, if it is determined that the instruction information has been superimposed on the still image, in step S198, the control unit 22G starts recording the video, first audio, second audio, and still image received by the communication unit 21G into the memory 23G.

[0388] In this way, when a support worker in a remote location extracts still images to be used to support the work from the video displayed on the display unit 34G of the support terminal 3G, and superimposes instruction information onto the still images, recording of the video and the first audio to memory 23G begins. Therefore, the worker does not need to perform any operation to start recording; the support worker in a remote location can record the video and the first audio to memory 23G by extracting still images from the video and superimposing instruction information onto the still images.

[0389] Furthermore, in this embodiment 8, the support terminal 3G may, instead of transmitting still images extracted from the video, transmit drawing data of the work target created with CAD (Computer Aided Design), operation manual data showing how to operate the work target, or an image of the entire screen displayed on the display unit 34G to the work terminal 1G via the server 2G. The control unit 22G may, triggered by receiving drawing data of the work target, operation manual data of the work target, or an image of the entire screen displayed on the display unit 34G of the support terminal 3G from the support terminal 3G, start recording the video and the first audio to the memory 23G. The work terminal 1G may display drawing data of the work target, operation manual data of the work target, or an image of the entire screen displayed on the display unit 34G of the support terminal 3G.

[0390] (Embodiment 9) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 9, however, recording of the video and first audio to memory is initiated when mode information received from the work terminal indicates a second mode in which the support worker uses the work terminal, and a predetermined keyword stored in advance is included in the first audio, which includes the voice of the support worker at the work site.

[0391] Figure 29 is a diagram showing the configuration of the work support system 10J according to this embodiment 9.

[0392] The work support system 10J shown in Figure 29 comprises a work terminal 1J and a server 2J. In this embodiment 9, the same reference numerals are used for components that are the same as in embodiment 1, and their descriptions are omitted.

[0393] In Embodiment 9, the worker performing the work is at the work site, and the support worker assisting the worker is also at the work site. At the work site, the support worker uses the work terminal 1J to assist the work while photographing the work object. In Embodiment 9, there are cases where the worker uses the work terminal 1J and cases where the support worker uses the work terminal 1J.

[0394] The work terminal 1J includes a communication unit 11J, a control unit 12, a memory 13, an input unit 14J, a camera 15, and a microphone 16. The following description of the work terminal 1J will explain the differences from the work terminal 1 of Embodiment 1.

[0395] The input unit 14J includes a switch to switch between a first mode in which an operator uses the work terminal 1J and a second mode in which a support worker uses the work terminal 1J. An operator using the work terminal 1J switches the switch to the first mode, and a support worker using the work terminal 1J switches the switch to the second mode.

[0396] The communication unit 11J transmits to the server 2J the video image acquired by the camera 15, the first audio acquired by the microphone 16, and mode information indicating whether either the first mode or the second mode received by the input unit 14J has been selected.

[0397] Server 2J comprises a communication unit 21J, a control unit 22J, and a memory 23J. Server 2J is an example of an information processing device. The following description of Server 2J will explain the differences from Server 2 of Embodiment 1.

[0398] The communication unit 21J receives video, first audio, and mode information from the work terminal 1J.

[0399] Memory 23J stores a predetermined start keyword and a predetermined end keyword in advance. The predetermined start keyword is, for example, a demonstrative pronoun such as "there" or "over there," the name of the work object, or the name of the work object's part. The predetermined end keyword is, for example, a phrase to end recording, such as "end recording." The predetermined start keyword and predetermined end keyword may be entered by a supporter. Memory 23J may store one start keyword or multiple start keywords. Similarly, memory 23J may store one end keyword or multiple end keywords.

[0400] The control unit 22J starts recording the video and first audio to memory 23J when the mode information received by the communication unit 21J indicates the second mode and a predetermined start keyword pre-stored in memory 23J is included in the first audio. The control unit 22J also stops recording the video and first audio to memory 23J when the mode information received by the communication unit 21J indicates the second mode and a predetermined end keyword pre-stored in memory 23J is included in the first audio.

[0401] More specifically, the control unit 22J determines whether the mode information received by the communication unit 21J indicates the first mode or the second mode. If it determines that the mode information indicates the second mode, the control unit 22J performs speech recognition on the first audio received by the communication unit 21J and converts the first audio into text. The control unit 22J then determines whether a predetermined start keyword, which is pre-stored in the memory 23J, is included in the converted first audio. If it determines that the predetermined start keyword is included in the first audio, the control unit 22J starts recording the video and the first audio into the memory 23J.

[0402] For example, an assistant wearing a work terminal 1J will converse with a worker while assisting with work at a work site. The control unit 22J determines that the assistant has started assisting with work if a predetermined start keyword is included in the first audio collected at the work site, and starts recording the video and the first audio to the memory 23J.

[0403] Furthermore, if it is determined that a predetermined start keyword is not included in the first audio, the control unit 22J determines whether a predetermined end keyword, which is pre-stored in memory 23J, is included in the text-converted first audio. If it is determined that the predetermined end keyword is included in the first audio, the control unit 22J terminates recording the video and the first audio to memory 23J.

[0404] Next, the work support processing performed by the work terminal 1J and the server 2J in Embodiment 9 of this disclosure will be described.

[0405] Figure 30 is a flowchart illustrating the work support processing performed by the work terminal 1J in Embodiment 9 of this disclosure.

[0406] First, in step S251, the input unit 14J accepts a selection from either the worker or the supporter, between a first mode in which the worker uses the work terminal 1J, and a second mode in which the supporter uses the work terminal 1J. The worker using the work terminal 1J switches the input unit 14J to the first mode, and the supporter using the work terminal 1J switches the input unit 14J to the second mode.

[0407] The processes in steps S252 to S253 are the same as those in steps S1 to S2 shown in Figure 2, so their explanation will be omitted.

[0408] Next, in step S254, the communication unit 11J transmits the video footage acquired by the camera 15, the first audio recording acquired by the microphone 16, and mode information indicating either the first or second mode received by the input unit 14J to the server 2J. At this time, the input unit 14J accepts input operations from an assistant or operator to start transmitting the video footage, the first audio recording, and the mode information.

[0409] Next, in step S255, the control unit 12 determines whether or not to terminate the transmission of the video, first audio, and mode information. At this time, the input unit 14J accepts an input operation from an assistant or operator to terminate the transmission of the video, first audio, and mode information. If an input operation to terminate the transmission of the video, first audio, and mode information is accepted, the control unit 12 determines to terminate the transmission of the video, first audio, and mode information. If an input operation to terminate the transmission of the video, first audio, and mode information is not accepted, the control unit 12 determines not to terminate the transmission of the video, first audio, and mode information.

[0410] If it is determined at this point that the transmission of the video, first audio, and mode information has ended (YES in step S255), the work support process ends. At this time, the communication unit 11J ends the transmission of the video, first audio, and mode information.

[0411] On the other hand, if it is determined that the transmission of the video, first audio, and mode information is not to be terminated (NO in step S255), the process returns to step S252.

[0412] Figure 31 is a flowchart illustrating the work support processing by server 2J in Embodiment 9 of this disclosure.

[0413] First, in step S261, the communication unit 21J receives the video, first audio, and mode information transmitted by the work terminal 1J.

[0414] Next, in step S262, the control unit 22J determines whether the mode information received by the communication unit 21J indicates a second mode.

[0415] If it is determined that the mode information does not indicate the second mode, that is, if it is determined that the mode information indicates the first mode (NO in step S262), the process returns to step S261.

[0416] On the other hand, if it is determined that the mode information indicates the second mode (YES in step S262), in step S263, the control unit 22J performs speech recognition on the first audio received by the communication unit 21J and converts the first audio into text.

[0417] Note that the processing in steps S264 to S268 is the same as the processing in steps S173 to S177 shown in Figure 22, so the explanation will be omitted.

[0418] In this way, when a support worker assisting with work using the work terminal 1J at the work site utters a predetermined keyword that has been stored in advance, the recording of the video and first audio to memory 23J begins. Therefore, the worker does not need to perform any operation to start recording, and the video and first audio can be recorded to memory 23J simply by a support worker assisting with work using the work terminal 1J at the work site uttering a predetermined keyword.

[0419] In this embodiment 9, the control unit 22J determines whether a predetermined termination keyword is included in the first audio, but this disclosure is not limited thereto. The work terminal 1J may accept a recording termination button being pressed by the supporter. If the recording termination button is pressed by the supporter, the work terminal 1J may send a recording termination signal to the server 2J. The communication unit 21J of the server 2J may receive the recording termination signal sent by the work terminal 1J. The control unit 22J may determine whether the recording termination signal sent by the work terminal 1J has been received. If it is determined that the recording termination signal has been received, the control unit 22J may terminate the recording of the video and the first audio to the memory 23J.

[0420] Furthermore, in this embodiment 9, the server 2J determines whether the mode information indicates a second mode and whether a predetermined start keyword is included in the first audio, but this disclosure is not particularly limited thereto. The work terminal 1J may also determine whether the mode information indicates a second mode and whether a predetermined start keyword is included in the first audio. If it is determined that the mode information indicates a second mode and that a predetermined start keyword is included in the first audio, the work terminal 1J may send a recording start signal to the server 2J to instruct the start of recording. When the server 2J receives the recording start signal sent by the work terminal 1J, it may start recording the video and the first audio to the memory 23J. The work terminal 1J may also determine whether the mode information indicates a second mode and whether a predetermined end keyword is included in the first audio. If it is determined that the mode information indicates a second mode and that a predetermined end keyword is included in the first audio, the work terminal 1J may send a recording end signal to the server 2J to instruct the end of recording. When server 2J receives a recording end signal transmitted by work terminal 1J, it may terminate the recording of the video and first audio to memory 23J.

[0421] (Embodiment 10) In Embodiment 1, recording of the video and first audio to memory 23 is initiated when a recording start signal is received. In Embodiment 10, recording of the video and first audio to memory is initiated when mode information received from the work terminal indicates a second mode in which the supporter uses the work terminal, and when the same object is continuously captured in a predetermined area of ​​the video for a predetermined time or longer.

[0422] Figure 32 shows the configuration of the work support system 10K according to this embodiment 10.

[0423] The work support system 10K shown in Figure 32 includes a work terminal 1J and a server 2K. In this embodiment 10, the same reference numerals are used for components that are the same as those in embodiments 1 and 9, and their descriptions are omitted.

[0424] In Embodiment 10, the worker performing the work is at the work site, and the support worker assisting the worker is also at the work site. At the work site, the support worker uses the work terminal 1J to photograph the work object while assisting the worker. In Embodiment 10, there are cases where the worker uses the work terminal 1J and cases where the support worker uses the work terminal 1J.

[0425] Server 2K comprises a communication unit 21K, a control unit 22K, and memory 23K. Server 2K is an example of an information processing device. The following description of Server 2K will explain the differences from Server 2 of Embodiment 1.

[0426] The communication unit 21K receives video, first audio, and mode information from the work terminal 1J.

[0427] The control unit 22K starts recording the video and first audio to memory 23K when the mode information received by the communication unit 21K indicates the second mode and the same object is continuously depicted in a predetermined area of ​​the video for a predetermined time or longer, which acts as a trigger. After recording of the video and first audio to memory 23K has started, the control unit 22K stops recording the video and first audio to memory 23K when the mode information received by the communication unit 21K indicates the second mode and the same object is not continuously depicted in a predetermined area of ​​the video for a predetermined time or longer, which acts as a trigger.

[0428] More specifically, the control unit 22K determines whether the mode information received by the communication unit 21K indicates the first mode or the second mode. If it determines that the mode information indicates the second mode, the control unit 22K analyzes the video image received by the communication unit 21K and determines whether the same object is continuously depicted in a predetermined area of ​​the video image for a predetermined time or longer. The predetermined area is the area that includes the center of each of the multiple still images that make up the video image. If it is determined that the same object is continuously depicted in the predetermined area of ​​the video image for a predetermined time or longer, the control unit 22K starts recording the video image and the first audio to the memory 23K.

[0429] For example, when an assistant wearing a work terminal 1J assists with work at a work site, they gaze at the work object. At this time, the same object (work object) is continuously captured in a predetermined area of ​​the video for a predetermined period of time or longer. When the assistant gazes at the work object at the work site, the control unit 22K determines that the assistant has started assisting with the work and begins recording the video and the first audio to memory 23K.

[0430] Furthermore, if, after recording of the video and first audio to memory 23K has started, it is determined that the same object has not been continuously captured in a predetermined area of ​​the video for a predetermined time or longer, the control unit 22K terminates recording of the video and first audio to memory 23K.

[0431] Next, we will describe the work support processing by Server 2K in Embodiment 10 of this disclosure.

[0432] Figure 33 is a flowchart illustrating the work support processing by Server 2K in Embodiment 10 of this disclosure.

[0433] The processes in steps S271 to S272 are the same as those in steps S261 to S262 shown in Figure 31, so their explanation will be omitted.

[0434] Next, in step S273, the control unit 22K analyzes the video image received by the communication unit 21K. The control unit 22K recognizes objects that are captured in a predetermined region including the center of each of the multiple still images that make up the video image.

[0435] Next, in step S274, the control unit 22K determines whether the same object is continuously captured in a predetermined area of ​​the moving image for a predetermined time or longer.

[0436] If it is determined that the same object is continuously captured in a predetermined area of ​​the video for a predetermined time or longer (YES in step S274), then in step S275, the control unit 22K starts recording the video and the first audio to the memory 23K. After that, the process returns to step S271.

[0437] On the other hand, if it is determined that the same object is not continuously captured in a predetermined area of ​​the video for a predetermined time or longer (NO in step S274), in step S276, the control unit 22K determines whether or not the video and the first audio are being recorded.

[0438] If it is determined that the video and first audio are not being recorded (NO in step S276), the process returns to step S271.

[0439] On the other hand, if it is determined that the video and first audio are being recorded (YES in step S276), in step S277, the control unit 22K terminates recording the video and first audio to memory 23K. Then, the process returns to step S271.

[0440] Thus, when a support worker using the work terminal 1J to assist with work at the work site gazes at the work object, the same object will appear continuously in a predetermined area of ​​the video for a predetermined time or longer. Therefore, when a support worker uses the work terminal 1J at the work site and the same object appears continuously in a predetermined area of ​​the video for a predetermined time or longer, recording of the video and the first audio to memory 23K begins. Consequently, the worker does not need to perform any operation to start recording; the video and the first audio can be recorded to memory 23K simply by a support worker using the work terminal 1J to assist with work at the work site gazing at the work object.

[0441] In this embodiment 10, the server 2K determines whether the mode information indicates the second mode and whether the same object is continuously depicted in a predetermined area of ​​the video for a predetermined time or longer, but this disclosure is not limited thereto. The work terminal 1J may also determine whether the mode information indicates the second mode and whether the same object is continuously depicted in a predetermined area of ​​the video for a predetermined time or longer. If it is determined that the mode information indicates the second mode and that the same object is continuously depicted in a predetermined area of ​​the video for a predetermined time or longer, the work terminal 1J may send a recording start signal to the server 2K to instruct the start of recording. When the server 2K receives the recording start signal sent by the work terminal 1J, it may start recording the video and the first audio to memory 23K. Subsequently, if it is determined that the mode information indicates the second mode and that the same object is not continuously depicted in a predetermined area of ​​the video for a predetermined time or longer, the work terminal 1J may send a recording end signal to the server 2K to instruct the end of recording. When server 2K receives a recording end signal transmitted by work terminal 1J, it may terminate the recording of the video and first audio to memory 23K.

[0442] In embodiments 1 to 10, the support terminal or work terminal may accept input from the supporter or worker regarding the content of the video and audio recorded on the server.

[0443] Figure 34 shows an example of a screen displayed on the display unit 34 of the support terminal in embodiments 1 to 10.

[0444] The display unit 34 plays back the video and audio recorded on the server, and also displays a display screen 360 for receiving input from a supporter regarding the content of the video and audio.

[0445] The display screen 360 includes a search criteria input field 361 for searching for files, a file selection field 362 for accepting file selections, a playback field 363 for playing the selected file, and an information input field 364 for accepting input of information about the file content.

[0446] The video and audio are recorded as a single file. The server's memory records the date and time of the operation, username, equipment ID, call memo, event memo, and file name, associating them with the file.

[0447] The work date and time indicate the date and time when the video and first audio were recorded. The username indicates the name of the supporter or worker. The equipment ID indicates identification information to identify the equipment on which the work was performed. The call memo and event memo indicate information about the contents of the file. The file name indicates the name of the file.

[0448] The supporter enters at least one of the following into the search criteria input field 361: work date and time, username, equipment ID, call memo, event memo, and file name. This displays information about files that match the criteria entered in the search criteria input field 361 in the file selection field 362. If no criteria are entered in the search criteria input field 361, information about multiple recorded files is displayed in the file selection field 362. The file selection field 362 displays the work date and time, username, equipment ID, call memo, event memo, file name, play button, download button, and delete button.

[0449] When the play button is pressed, the corresponding file is played in the playback field 363. When the download button is pressed, the corresponding file is downloaded from the server to the support terminal. When the delete button is pressed, the corresponding file is deleted from the server's memory. The information input field 364 accepts input of call memos and event memos for the file played in the playback field 363.

[0450] In Figure 34, a support terminal displays the display screen 360 and accepts information input from a support worker, but this disclosure is not limited to this. A work terminal may display the display screen 360 and accept information input from a worker, or terminals other than the support terminal and work terminal may display the display screen 360 and accept information input from a support worker or a worker.

[0451] In each of the above embodiments, each component may be implemented by dedicated hardware or by executing a software program suitable for each component. Each component may also be implemented by a program execution unit such as a CPU or processor reading and executing a software program recorded on a recording medium such as a hard disk or semiconductor memory. Furthermore, the program may be executed by another independent computer system by recording and transferring the program to a recording medium, or by transferring the program via a network.

[0452] Some or all of the functions of the apparatus according to the embodiments of this disclosure are typically implemented as an integrated circuit, or LSI (Large Scale Integration). These may be individually integrated onto a single chip, or some or all of them may be integrated onto a single chip. Furthermore, the integration is not limited to LSIs, but may also be implemented using dedicated circuits or general-purpose processors. An FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells inside the LSI may also be used.

[0453] Furthermore, some or all of the functions of the apparatus according to the embodiments of this disclosure may be realized by a processor such as a CPU executing a program.

[0454] Furthermore, all figures used above are illustrative examples provided to illustrate this disclosure, and this disclosure is not limited to these illustrative figures.

[0455] Furthermore, the order in which the steps shown in the flowchart above are performed is illustrative for the purpose of specifically illustrating this disclosure, and other orders may be used to the extent that similar effects can be achieved. Also, some of the above steps may be performed simultaneously (in parallel) with other steps. [Industrial applicability]

[0456] The technology disclosed herein can reduce the amount of data to be recorded in memory and alleviate the burden on workers, making it useful as a technology for recording video footage and audio collected at work sites on a server. [Explanation of symbols]

[0457] 1,1A,1B,1E,1F,1G,1J Work terminals 2,2A,2B,2C,2D,2E,2F,2G,2J,2K Server 3,3A,3B,3C,3G Support terminal 4 Network 10, 10A, 10B, 10C, 10D, 10E, 10F, 10G, 10J, 10K Work Support System 11,11A,11B,11E,11F,11G,11J Communication department 12, 12E, 12G Control Unit 13 memory 14, 14E, 14J Input Section 15 Cameras 16 Microphones 17 speakers 18 GPS receiver 19 Beacon receiver 20 Display unit 21, 21A, 21B, 21C, 21D, 21E, 21F, 21G, 21J, 21K Communication unit 22, 22A, 22B, 22C, 22D, 22E, 22F, 22G, 22J, 22K Control unit 23, 23A, 23B, 23C, 23E, 23F, 23G, 23J, 23K Memory 31, 31A, 31B, 31C, 31G Communication unit 32, 32C, 32G Control unit 33 Memory 34, 34A, 34B, 34G Display unit 35 Speaker 36 Microphone 37, 37A, 37B, 37C, 37G Input unit

Claims

1. A method of information processing performed by a computer, Receiving video footage and audio recordings taken at the work site from a work terminal used by a worker at the work site, The recording of the video and the first audio to memory is initiated by a support worker assisting the aforementioned worker with the task, Information processing methods including

2. Furthermore, the system includes receiving a recording start signal from a support terminal used by the supporter, which instructs the start of the recording based on an input operation by the supporter located remotely. The commencement of the recording includes, triggered by the receipt of the recording commencement signal, initiating the recording of the video and the first audio to the memory. The information processing method according to claim 1.

3. Furthermore, area designation information indicating a predetermined area at the work site, as specified by the supporter located in a remote location, is received from the support terminal used by the supporter. Furthermore, the system receives location information from the work terminal indicating the current location of the work terminal, Includes, The commencement of the recording includes starting the recording of the video and the first audio to the memory, triggered when the location of the work terminal indicated by the location information enters the predetermined area indicated by the area designation information. The information processing method according to claim 1.

4. Furthermore, the support terminal used by the supporter receives identification information contained in the wireless signal transmitted from the work object at the work site designated by the supporter located in a remote location. Furthermore, the work terminal receives from the work terminal the identification information contained in the wireless signal received by the work terminal and signal information including the radio wave intensity of the wireless signal measured by the work terminal, Includes, The commencement of the recording includes initiating the recording of the video and the first audio to the memory, triggered by the radio wave intensity of the wireless signal received from the work terminal, which includes the identification information received from the support terminal, being above a threshold value. The information processing method according to claim 1.

5. Furthermore, the system includes receiving a second audio signal from the support terminal used by the supporter located in a remote location, The commencement of the recording includes initiating the recording of the video and the first audio to the memory when a predetermined keyword stored in advance is included in the second audio, The information processing method according to claim 1.

6. The commencement of the recording includes detecting the utterance section spoken by the supporter in the second audio, and triggering the recording of the video and the first audio to the memory when the detected utterance section contains a predetermined keyword that has been stored in advance. The information processing method according to claim 5.

7. The commencement of the recording includes recognizing the actions of the support worker at the work site from the received video, and triggering the recording of the video and the first audio to the memory when the recognized action is a predetermined action. The information processing method according to claim 1.

8. The commencement of the recording includes triggering the recording of the video and the first audio to the memory when a predetermined keyword stored in advance is included in the first audio, which includes the voice of the supporter present at the work site. The information processing method according to claim 1.

9. Furthermore, the received video and audio are transmitted to a support terminal used by the supporter located remotely. Furthermore, still images extracted by the supporter from the moving image displayed on the support terminal's display unit are received from the support terminal. Includes, The commencement of the recording includes, triggered by the reception of the still image, initiating the recording of the moving image and the first audio to the memory. The information processing method according to claim 1.

10. Furthermore, the system includes receiving a second audio signal from the support terminal, The commencement of the recording includes, triggered by the reception of the still image, initiating the recording of the moving image, the first audio, the second audio, and the still image to the memory. The information processing method according to claim 9.

11. Furthermore, the received video and audio are transmitted to a support terminal used by the supporter located remotely. Furthermore, still images extracted by the supporter from the moving image displayed on the support terminal's display unit are received from the support terminal. Includes, The reception of the still image includes receiving the still image from the support terminal, on which instruction information entered by the supporter is superimposed using the support terminal. The commencement of the recording includes triggering the superimposition of the instruction information onto the still image to start recording the moving image and the first audio to the memory, The information processing method according to claim 1.

12. Furthermore, the system includes receiving mode information from the work terminal indicating whether a first mode in which the worker uses the work terminal or a second mode in which the supporter uses the work terminal has been selected. The commencement of the recording includes initiating the recording of the video and the first audio to the memory, triggered by the received mode information indicating the second mode and the inclusion of a predetermined keyword stored in advance in the first audio. The information processing method according to claim 1.

13. Furthermore, the system includes receiving mode information from the work terminal indicating whether a first mode in which the worker uses the work terminal or a second mode in which the supporter uses the work terminal has been selected. The commencement of the recording includes initiating the recording of the video and the first audio to the memory, triggered by the received mode information indicating the second mode and the same object being continuously captured in a predetermined area of ​​the video for a predetermined time or longer. The information processing method according to claim 1.

14. Communications Department and, Control unit and Memory and Equipped with, The communication unit receives video footage and audio recordings taken at the work site from a work terminal used by a worker at the work site. The control unit, triggered by a recording start operation performed by an assistant assisting the worker, starts recording the video and the first audio to the memory. Information processing device.

15. The video footage captured at the work site and the first audio recording collected at the work site are received from a work terminal used by a worker at the work site. The computer is configured to start recording the video and the first audio to memory, triggered by a recording start operation performed by a supporter assisting the worker. Information processing program.

Citation Information

Patent Citations

  • Remote assistance system for person requiring assistance

    JP2001344355A

  • Work remote support system and work remote support method

    JP2015118714A