Video quality feedback and deficiency identification for video risk assessement

The method and system improve video quality for risk assessment by identifying and correcting deficiencies in real-time, enhancing the effectiveness of video recordings for hazard detection and worker safety.

WO2025141281A1PCT designated stage expired Publication Date: 2025-07-03FYLD LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/GB2023/053382
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

Existing video risk assessment methods in industries like construction and energy suffer from poor-quality video recordings due to issues such as camera orientation, audio gaps, and insufficient narration, which compromise the effectiveness of risk assessment.

Method used

A method and system that analyze video recordings in real-time using a mobile device or remote server to identify deficiencies, provide prompts for correction, and ensure high-quality video capture by field workers through visual and audible feedback.

Benefits of technology

Enhances the quality of video recordings for risk assessment, improving site visibility and worker safety by ensuring comprehensive hazard identification and reducing the need for re-recording.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2023053382_03072025_PF_FP_ABST
    Figure GB2023053382_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A method of machine-analysed video risk assessment is presented. The method comprising: analysing a video recorded using a video recorder of a mobile device (102) at a location where a job is being performed; identifying whether a deficiency is present in the video that would lead to a compromised video risk assessment; issuing a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency; and analysing the video after the prompt has been issued to identify whether a deficiency is present in the video that would lead to a compromised video risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

VIDEO QUALITY FEEDBACK AND DEFICIENCY IDENTIFICATION FOR VIDEO RISK ASSESSEMENT

[0001] The present disclosure relates generally to the field of video risk assessment, and more particularly to a method of machine-analysed video risk assessment and a related video risk assessment system.BACKGROUND

[0002] In general, industries such as, but not limited to, construction industries, energy and utilities industries, or the like, involve day-to-day work (also be referred to as field work, field operations, or the like) to be performed at a location / work site. The work at the location has to be performed in compliance with field work management requirements defined based on state, local, and federal laws, a type of work, a work site, and so on. The field work management requirements include multiple regulations, codes and standards with regard to health and safety of workers, environmental protection, and quality management.

[0003] Performing risk assessment before, during and after a job may be encouraged or required in these types of industries. Video risk assessment can be used to identify risks in an area such as a jobsite, where work needs to be done efficiently without accidents or inconveniences slowing the work down or even stopping the work. Any field worker on the site can use a video risk assessment, VRA, application on a mobile phone or other remote device that includes a camera, capturing video of the area in which they intend to work for video analysis.

[0004] The video analysis for risk assessment requires a certain level of quality in a video to perform proper analysis. A VRA video may be of poor quality or otherwise unusable (for example, failing to capture the jobsite or enough of the jobsite or the right part of the jobsite) due to the camera being pointed at the ground for an extended period of time, a gap in or lack of voiceover describing the area, moving the camera around too quickly or too slowly and the like. Provided that the VRA video is of good quality, the area can be automatically assessed quickly by the model. Therefore, prompting the field worker to produce a good quality video would be beneficial.

[0005] Identifying deficiencies in videos such that the videos can be retaken may also be useful in other video capturing areas such as driving hazard perception videos, street-view recordings for informational purposes such as mapping and directions, and even entertainment purposes (for example, when filming for movies or the like).

[0006] It is an aim of the present disclosure to improve video capture for analysis.SUMMARY STATEMENTS

[0007] The following statements summarise some aspects and embodiments of the disclosed technology, however, the scope of the invention is as defined by the accompanying claims.

[0008] According to a first aspect a method of machine-analysed video risk assessment is provided. The method comprising the following steps: (i) analysing a video recorded using a video recorder of a mobile device at a location where a job is being performed; (ii) identifying whether a deficiency is present in the video that would lead to a compromised video risk assessment; (iii) issuing a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency; and (iv) analysing the videoafter the prompt has been issued to identify whether a deficiency is present in the video that would lead to a compromised video risk assessment.

[0009] The provided method provides for high quality risk assessment video which ensures that fieldworker records the jobsite and narrates the recording as expected. This will help in reducing the risk of incident and positively impacting fieldworker safety. For the remote video high quality risk assessment video provides better site visibility and fuller details on potential risks, better informing their decision on the need of intervention

[0010] The method may comprise (v) issuing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed.

[0011] The method may comprise: at step (iii) instead of issuing a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency, issuing a first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; and at step (v) instead of issuing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed, issuing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies at a time after the first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies.

[0012] The steps (i) to (iii) may be performed while the video is being recorded.

[0013] The step (iv) may be performed while the video is being recorded.

[0014] The steps (i) and (iv) may be performed at the mobile device.

[0015] The steps (i) and (iv) may be performed remote from the mobile device at a remote server.

[0016] The remote server may be a cloud server.

[0017] A field of view of the video being recorded may be displayed on a screen of the mobile device for a user to view while the video is being recorded and wherein the visual or audible prompt is a message displayed on the screen of the mobile device while the video is being displayed.

[0018] At least one of steps (i), (ii) and (iv) may be performed using an artificial intelligence model trained to identify one or more types of deficiency in a risk assessment video.

[0019] At least one of steps (i), (ii) and (iv) may be performed using at least one heuristic marker to identify the presence or absence of a deficiency.

[0020] The visual or audible prompt may comprise a quality indicator of the video. The quality indicator may be a message recommending that the user produces a new video.

[0021] According to a second aspect a video risk assessment system is provided. The video risk assessment system comprising: a video recorder; one or more processors; and one or more memories comprising program code instructions for implementing the method according to the first aspect, when executed on the one or more processors.

[0022] According to a third aspect a non-transitory computer-readable storage medium is provided.The non-transitory computer-readable storage medium having stored thereon instructions forimplementing the method according to the first aspect, when executed on one or more devices having processing capabilities.

[0023] According to a fourth aspect a computer-implemented method of providing training to a field worker to improve safety on a jobsite is provided. The method comprising: analysing a video recorded using a video recorder of a mobile device at a location where a job is being performed, wherein analysing the video comprises identifying one or more quality markers and wherein the mobile device is associated with a field worker involved in the job; identifying whether a deficiency is present in the video that would lead to a compromised video risk assessment, based on the one or more quality markers; generating a risk assessment quality score from the video, based on the one or more quality markers; providing historical risk assessment data labelled with historical quality markers from at least one earlier video recorded by the same field worker; inputting the risk assessment quality score and historical risk assessment data into a risk assessment model configured to generate an up-to-date risk assessment quality score for the field worker and running the model to create an up-to-date risk assessment quality score for the field worker; and providing an indication of safety performance, based on the up-to-date risk assessment quality score, and enabling training content to be transmitted to the mobile device and / or a user in a supervisory role based on the indication of safety performance.

[0024] The computer-implemented method may further comprise: providing a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency, or providing a first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; analysing the video after the prompt has been provided to identify whether a deficiency is present in the video that would lead to a compromised video risk assessment; providing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed or at a time after the first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; and updating the field worker's up-to- date risk assessment quality score based on one or both of the first and second visual or audible prompts.

[0025] The one or more quality markers may be identified using an Al model trained to analyse a video and identify the one or more quality markers in image and / or audio information of the video.

[0026] The indication of safety performance may be provided to the field worker as a visual or audible prompt at the mobile device.

[0027] The computer-implemented method may comprise: at a remote server, identifying the up-to- date risk assessment quality score as belonging to a field worker in a defined team of field workers; and at the remote server, automatically aggregating the up-to-date risk assessment quality score with at least one risk assessment quality score of another field worker belonging to the same defined team of field workers to create an up-to-date team risk assessment quality score.

[0028] According to a fifth aspect a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium having stored thereon instructions for implementing the computer-implemented method according to the fourth aspect, when executed on one or more devices having processing capabilities.BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The foregoing will be apparent from the following more particular description of the example embodiments, as illustrated in the accompanying drawings in which like reference characters refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the example embodiments.

[0030] Fig. 1 illustrate a fieldworker interacting with a mobile computing device of a risk assessment system at a jobsite comprising a plurality of hazards.

[0031] Fig. 2 is a block diagram of a method of machine-analysed video risk assessment.

[0032] Fig. 3 illustrate a risk assessment system and a fieldworker interacting with a mobile computing device of the risk assessment system at a jobsite comprising a plurality of hazards.

[0033] Figs 4a-d illustrate screenshots of different implementations of different steps in the method of Fig. 2.

[0034] Fig. 5 is an example data flow for video analysis.

[0035] Fig. 6 shows an example mobile device.

[0036] Fig. 7 shows a video, represented by a sequence of frames and audio.

[0037] Fig. 8 is a block diagram of a computer-implemented method of providing training to a field worker to improve safety on a jobsite.DETAILED DESCRIPTION

[0038] Aspects of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. The methods and apparatus disclosed herein can, however, be realized in many different forms and should not be construed as being limited to the aspects set forth herein. Like numbers in the drawings refer to like elements throughout.

[0039] The terminology used herein is for the purpose of describing particular aspects of the disclosure only, and is not intended to limit the invention. It should be emphasized that the term "comprises / comprising" when used in this specification is taken to specify the presence of stated features, integers, steps, or components, but does not preclude the presence or addition of one or more other features, integers, steps, components, or groups thereof. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0040] Embodiments of the present disclosure will be described and exemplified more fully hereinafter with reference to the accompanying drawings. The solutions disclosed herein can, however, be realized in many different forms and should not be construed as being limited to the examples set forth herein.

[0041] It will be appreciated that when the present disclosure is described in terms of a method, it may also be embodied in one or more processors and one or more memories coupled to the one or more processors, wherein the one or more memories store one or more programs that perform the steps, services and functions disclosed herein when executed by the one or more processors.

[0042] In the following description of exemplary embodiments, the same reference numerals denote the same or similar components.

[0043] Before the methods of the present disclosure can analyse a video, a field worker 100 needs to capture the video of the jobsite 120. The field worker 100 captures a video using a mobile device 102. A video comprises visual information in a series of frames and may additionally comprise audio information. Risks may be assessed based on both the visual and audio contents in a video. A video to be used for VRA may be around 1 minute, 1.5 minutes, 2 minutes in duration or another suitable duration. The duration of the video for VRA may be sufficiently long for the field worker 100 to indicate the hazards that they can see and any other notable things that they can see or know about the jobsite. Analysis of the video for any deficiencies may be performed on a "clip" of the video (i.e. not the whole video), meaning analysis needs to be quick in order to provide feedback to the field worker 100 within the time it takes to capture a video of a duration of around 1 minute, for example.

[0044] Figure 1 shows an example field worker 100, who will capture a video at a jobsite 120 in order for a video risk assessment to be performed. Figure 1 depicts an example mobile device 102 comprising a video module 104. The video module 104 may comprise a video recorder configured to record video. In this example, a lens of the video module 104 is shown on the back of the mobile device 102; Figure 1 also shows the front of the mobile device 102 which in this example comprises a display 106. The mobile device 102 may additionally or alternatively include a lens of the video module 104 on the front side - for example, to allow a field worker 100 to capture video including their face using a front-facing camera.

[0045] The mobile device 102 may comprise a microphone 110 configured to record audio. The video recorder may comprise both a camera to capture visual information and a microphone to capture audio information. The video recorder may comprise the microphone 110 and the camera.

[0046] The mobile device may comprise a speaker 110a configured to play audio.

[0047] The mobile device 102 comprises a processor 108 and memory 112. The memory 112 comprises an operating system configured to operate the mobile device 102. The memory 112 may store a mobile application for video analysis according to the present disclosure.

[0048] The mobile device 102 may be a mobile phone or a tablet or laptop, for example, or another suitable mobile computing device. A computing device that the field worker 100 cannot carry around at a jobsite 120 for capturing video would not be suitable for the present purpose. A computing device moveable on wheels or the like may be used but may not be the most elegant solution to capturing video of a jobsite versus a handheld mobile device. The mobile device 102 may be a digital camera equipped with a processor 108 and memory 112 such that a mobile application can be run on the digital camera.

[0049] The video module 104, the microphone 110 and the speaker 110a may be operatively connected to the memory 112 and the processor 108. The mobile device 102 may further comprise controlling circuitry 114 configured to control operations of the mobile device 102. The mobile device 102 may comprise a transmitter and receiver or transceiver 116 configured to send and receive digital signals. The display 106 may comprise or be provided additionally to an I / O module 106a, which may comprise a keyboard or touchscreen to allow a user to input information into the mobile device 106. The I / O module 106a may comprise the display 106 and the speaker 110a, for example, such that information can be presented to the user asan output of analysis. The information may take the form of text, graphical representations such as charts or diagrams, or audio - for example a spoken message in a relevant language to the user or sounds such as bells or alarms or the like. The microphone 110 may be part of the I / O and may allow to field worker 100 to control operations of the mobile device 102 using voice commands, for example.

[0050] In Figure 1, the jobsite 120 is shown comprising three example hazards 122: cloudy weather (which may indicate that rain will fall at the jobsite or may reduce visibility), an area where entry is prohibited, a road and an electrical hazard, for example power lines. These are merely examples of hazards that could occur at a jobsite 120 that would affect how a job is carried out. The field worker 100 may capture a video of the jobsite 120 including one or more hazards. For video risk assessment purposes, the field worker 100 may spend time showing a hazard in detail (from a safe distance or in a safe manner) or may narrate their video to indicate that they have perceived the hazard - for example "I am showing a road that passes close to the jobsite. I have observed that the road is currently busy".

[0051] Figure 2 shows a method of machine-analysed video risk assessment 200. The method 200 is computer-implemented. At step (i) the video captured by the field worker 100 is analysed 201. In particular, the video is analysed to determine 202 whether any deficiencies are present in the video. Deficiencies include lack of video, lack of audio, poor quality video, poor quality audio, moving too fast or too slow, orienting the camera incorrectly, taking the video without a flash or other lighting when the jobsite is dark, capturing video of the ground instead of the jobsite and the like. Other deficiencies include the total duration of the video and the word count being too low, for example. Detail of the analysis follows herein.

[0052] If it is determined that the video is free from deficiencies, a first prompt may be issued 203 to the field worker 100 at the mobile device 102. A prompt may comprise a "quality indicator" based on the analysis of the video. The quality indicator may be an indication of a particular deficiency being present in the video or a general indication that a deficiency has been identified so action needs to be taken to improve the video. The quality indicator may include a particular deficiency such as "video too dark" or the like. The quality indicator may be accompanied by a suggested solution to the deficiency; for example, "use a flash" for a dark video.

[0053] In this example, the first prompt would positively indicate that the field worker 100 is doing a good job of capturing the video. For example, the first prompt may be a written message of encouragement - such as "good job" or "keep it up" or the like. The first prompt may be an audio message - such as "good job" or "keep it up" or the like. The audio prompt may be a sound having positive connotations such as a bell sound typically associated with a correct answer.

[0054] If it is determined that one or more deficiencies are present in the video, feedback needs to be issued to the field worker 100 to adjust how they are capturing the video. The first prompt issued 203a to the field worker 100 may indicate that a deficiency is present. For example, the first prompt may take any of the forms listed above but with a negative connotation. The message may read / recite that video is missing, that the audio is missing, that the video is too dark or too bright, that the camera appears to be upside down, that the microphone 110 appears to be covered and so the sound is muffled, that the camera is being moved too quickly and causing blurring and the like. The audio prompt may be a sound having positive connotations such as a bell sound typically associated with an incorrect answer.

[0055] The issued 203b first prompt may include a solution to guide the field worker 100 to remove the deficiency. For example, the audible or visual prompt may indicate that the field worker 100 needs to uncover the camera lens, uncover the microphone 110, keep the mobile device 102 steady, move slower or more quickly, tilt the mobile device 102, use a flash or other light, speak louder, speak more quickly or leave fewer gaps between speaking and the like.

[0056] The first prompt may include both an indication of the deficiency and a solution. For example, the first prompt may comprise a message (either visual or audible or both) indicating "Too dark. Please activate flash" or a similar combination of a deficiency and a proposed solution.

[0057] Once the field worker 100 has been delivered feedback in the form of the first prompt, the video may be analysed 204 at step (iv). Analysis may be performed automatically once the first prompt has been issued or may be triggered after a predetermined time from the first prompt, for example 5 seconds or 10 seconds. In this way, the field worker 100 has time to digest the prompt and adjust their video capturing before the analysis takes place. Analysis may continue from step 201 to step 204 and beyond - i.e. analysis may still take place while the first prompt is being issued.

[0058] Again, analysis of the video includes determining 204a whether any deficiencies are present in the video. The analysis may first determine whether the earlier deficiency is still present or has been removed, before performing further analysis to determine whether any other deficiencies are present.

[0059] If no deficiencies are identified in the video, a second prompt may be issued to the field worker 100 indicating that the video is free of deficiencies (for example in the manners described for the first prompt where no deficiencies are present).

[0060] If a deficiency is identified, a second prompt in line with the first negative prompt may be issued. The second prompt may merely indicate that a deficiency is present 205a, or may include a solution 205b. As shown, if there is a deficiency present, then the method proceeds to analysis 204 again and additional prompts up until the Nth prompt.

[0061] The method 200 comprise issuing such negative prompts until no deficiency are present in the video. Upon no deficiency are present in the video a positive final prompt may be issued. The final prompt indicates that the video is free of deficiencies. Additional prompts may be issued thereafter to encourage the field worker 100 to continue capturing a video without deficiencies - for example "keep it up" as described above. After a predetermined number of prompts indicating the presence of deficiencies and / or providing solutions, the method 200 may comprise indicating to the field worker 100 by way of a prompt that they should end the video. The prompt may indicate that the field worker 100 should restart the video capturing and then restart analysis, for example.

[0062] Some deficiencies may not render the video useless for VRA, but merely make the video less than ideal. As such, a prompt may indicate tips for the next time the field worker 100 captures a video, but does not suggest that the field worker 100 needs to recapture the video they have already taken.

[0063] The duration of a useful video for VRA may be 30 seconds, 60 seconds, 2 minutes or another duration. The video should be long enough that the field worker 100 can capture at least one hazard and indicate that they have perceived the hazard. The duration of a video that can beanalysed according to the present method may be 30 seconds, 60 seconds, 2 minutes or another duration. Analysis may not require a video duration as long as the VRA may require, it is merely important that the analysis can determine that the video will be useful for VRA. The time taken to identify and remove deficiencies may be less than the time required to capture all the hazards present at the jobsite.

[0064] Analysis may be performed on the video while the video is being recorded. In this way, the field worker 100 can adjust their video capturing based on the prompt(s) and improve the quality of the video for VRA.

[0065] At least steps (i) to (iii) of the method 200 of Figure 2 may be carried out during the video being captured. On the basis of the first prompt, the field worker 100 may pause and restart capturing the video having adjusted based on the first prompt, or the field worker 100 may stop capturing the first video and start again. The field worker 100 may continue capturing the video and make any adjustments based on the first prompt while the video is being captured.

[0066] Steps (iv) and / or (v) of the method 200 of Figure 2 may also be carried out during the video being captured. The field worker 100 may receive prompts during the video being captured to encourage them to continue correctly capturing the video or to adjust the video capturing so that the video is of good quality for VRA.

[0067] The mobile device 102 comprises a processor 108 as shown in Figure 1. The video analysis steps 201, 204 of the method 200 may be performed by the processor 108. Performing video analysis on the mobile device 102 may remove the need for a network connection in order to perform video analysis. Jobsites may be in low network coverage areas and it may be beneficial to remove the need for a network connection.

[0068] However, the video analysis steps 201, 204 may be performed at a remote location from the mobile device 102. The mobile device 102 may be part of a video analysis system 300 also comprising a remote server 302, as shown in Figure 3. The server 302 may comprise a processor 308 and memory 312 and may be configured to perform the analysis steps 201, 204. Any analysis step of the present method 200 may be performed at the server 302 and an output may be provided to the mobile device 102 for communication to the field worker 100. The mobile device 102 may be connectable to the server 302 using a communication network. In some examples, the communication network may include, but is not limited to, a wired network, a value-added network, a wireless network, a satellite network, or a combination thereof. Examples of the wired network may be, but is not limited to, a Local Area Network, LAN, a Wide Area Network, WAN, an Ethernet, and so on. Examples of the wireless network may be, but is not limited to, a cellular network, a wireless LAN, Wi-Fi, Bluetooth, Bluetooth low energy, Zigbee, Wi-Fi direct, WFD, Ultra-wideband, UWB, infrared data association, IrDA, near field communication, NFC, and so on. In some examples, the mobile device 102 may be connected to the server 302 directly (for example, via a direct communication, via an access point, or the like). In some examples, the mobile device 102 may be connected to the server 302 via a relay, a hub, and a gateway. The mobile device 102 may be connected to a cloud service provider using the communication network, or via the direct communication, or via the relay, the hub, and the gateway, or the like; the server 302 may be a cloud-based server.

[0069] Where the mobile device 102 may be connected to the server 302 via a wired connection, the server 302 may be located at the jobsite 120. As the field worker 100 may need to walkaround or otherwise capture different parts of the jobsite 120, a wireless connection to the server 302 may be more practical than a wired connection.

[0070] The mobile device 102 may be configured to transmit data to the server 302 using the transceiver 116. The controlling circuitry 114 may automatically cause the transceiver 116 to begin transmitting video data to the server 302 when the field worker 100 begins capturing the video.

[0071] Figures 4A and 4B show example user interfaces 400 as presented to a field worker 100 at the display 106. The mobile device 102 may comprise a user interface module (Ul module) configured to provide the user interface 400 through which input from the field worker 100 may be received and through which information may be displayed to the field worker 100 at the display 106, for example a view of what the video module 104 is picking up at the camera lens. The user interface 400 may form a part of the I / O module 106a because the field worker 100 is enabled to provide input via the user interface 400.

[0072] Figure 4A shows an example user interface 400 in which a prompt 402 is displayed adjacent the video feed being shown on the display 106. The display 106 may show the field of view of the video module 104 as a video feed on the display 106, to help the field worker understand what they are capturing - seeing the video feed will help the field worker aim the camera as desired and allow them to narrate the video based on what they can see. In this example, the prompt 402 is a written message and includes both an indication of the deficiency and a solution (turn on the flash). The user interface 400 has presented the field worker with an interactive element 404 (in this example, a touchscreen button labelled "OK"). When the field worker acknowledges that they have seen the prompt 402 by interacting with the interactive element 404, the prompt 402 may disappear, become minimised, or move to another position on the display 106 for example. Figure 4B shows another example user interface 400 where the prompt 402 has been overlaid onto the video feed shown on the display 106. This type of prompt 102 may also include an interactive element 404, for example an "X" or other commonly used icon that indicates that the prompt 402 can be closed, or the interactive element 404 may indicate that the prompt 402 can be moved on the display 106 to a position away from obscuring the video feed. The user interface 400 may enable the field worker to swipe the display 106 to remove or otherwise move the prompt 402 from / on the display 106. The controlling circuitry 114 may be configured to remove the prompt 402 or move the prompt 402 on the display 106 based on a voice command received at the microphone 110.

[0073] A prompt 402 may comprise an image. For example, an encouraging prompt 402 may comprise a happy face icon or a thumbs up icon or a tick / check or the like. Figure 4C shows an example of a prompt 402 comprising an image - a thumbs up icon. The chosen icon for an encouraging prompt 402 or a prompt indicating a deficiency (i.e. a negative prompt 402) may be selected based on the country of use. For example, a symbol widely understood to be positive or negative may mean something else in a specific country and so, to avoid confusion, the chosen image may have a universal meaning or may be country-specific.

[0074] Encouraging prompts 402 may include tips for producing a good quality video for VRA, such as "make sure to describe the context of the hazard" or the like, to draw out details from the field worker 100 that may not be identifiable using the Al model. For example, the model may be configured to determine whether the field worker 100 is speaking enough.

[0075] The image for use in the prompt 402 may represent the nature of the deficiency. A prompt 402 indicating that the video is too dark may comprise an icon resembling a light bulb or torch, for example.

[0076] An image may be simpler and / or faster for the field worker 100 to comprehend than a written message. A prompt 402 may comprise one or any combination of (i) text, (ii) image and (iii) sound. Figure 4D shows an example where the prompt 402 comprises only a sound. The field of view of the camera is not obscured, as shown.

[0077] The field worker 100 may be presented with an interactive prompt. The first and / or second and / or Nth prompt may comprise a mechanism for the field worker 100 to confirm that they have seen or heard the prompt. For example, the prompt may request that the user speaks into the microphone to acknowledge the prompt. A visual prompt appearing on the display 106 may comprise an interactive element (where the display 106 comprises a touchscreen) for the field worker 100 to acknowledge the prompt. Acknowledging the prompt may cause the prompt to disappear, or stop making a sound where the prompt is audible.

[0078] The display 106 may be configured to show a video feed representing the video being captured by the video module 104. When a visual prompt is presented to the field worker on the display 106, it may partially cover the video feed. Therefore, it may be desirable to remove or minimize or move the prompt after the prompt has been understood, so that the video feed is not obscured.

[0079] In addition to one or more prompts 402, the user interface 400 may comprise a clock such as a countdown clock to indicate to the field worker how long the video has been recording or how much longer a video needs to be for VRA. The user interface 400 may display user information such as field worker identification information, employer information and the like.

[0080] The user interface 400 may include one or more dynamic elements such as weather information or location information (for example, a live map). Having this information shown on the display 106 while capturing the video may inform the field worker's commentary, for example. The field worker may narrate their video (with their voice being captured at the microphone 110) and be able to identify that a hazard is present that is only considered a hazard at particular times of the day, in particular weather, at particular locations or the like. For example: "I see that it is due to rain. A pothole is present. Although the pothole is visible before rainfall, it may fill with water and become disguised as a puddle, meaning the pothole will be harder to spot." In another example: "I see a school next to the jobsite. Although there are no pedestrians or cars close-by now, I see that the time is 2:30pm and that parents and children will shortly be gathering very close to the jobsite, which is a potential hazard". The information provided at the user interface 400 that is not generated by the method 200 or captured by the video module 400 may be obtained from the Internet - for example a local news outlet or weather website - and / or may be obtained from information available to the mobile device 102 within one or more mobile applications such as a Weather Application, News Application, Maps Application or the like. The location information may be obtained using location tracking of the mobile device 102.

[0081] The user interface 400 may present training content to the field worker 100. For example, a prompt 402 may comprise a link to relevant training content, such as a training video on how to capture good videos for VRA or the like. A prompt 402 may contain training content. For example, where a deficiency is a lack of narrative and a solution is to narrate the video as it isbeing captured, training content may be provided in the prompt 402 as follows: "Too quiet. Please narrate the video. We suggest telling us what you can see and if you think it is a hazard now or may become a hazard." Providing training content while the video is being captured will improve video quality for VRA.

[0082] Video analysis may be performed after a video has been captured. The video may be stored in the memory 112 of the mobile device 102. The controlling circuitry 114 may be configured to cause the processor 108 to perform analysis of the video. Causing the processor 108 to begin analysis may occur automatically once capturing of the video is complete and the video has been stored. A prompt may be displayed in the user interface 400 even after the video has been captured; for example, the prompt may appear over an image of the final frame of the video, which may remain on the display 106, or the field of view of the camera may be replaced by a display of the prompt for example. According to another example, an interruption screen with the prompt may be issued. If the prompt indicates that there is a deficiency, which may include a solution, the prompt or another message on the user interface 400 may request that the field worker 100 begins capturing another video and tries to remove the deficiency in their next attempt.

[0083] Video analysis may be performed while the video is being captured, which may be referred to as "in-flight" video analysis. The controlling circuity 114 may be configured to cause the processor 108 to perform video analysis. This may be automatic once the field worker 100 has begun capturing the video.

[0084] It is beneficial to give feedback to the field worker during video capture to prevent them from having to repeat the same video to achieve a better quality. Repeating video capture could be a waste of valuable time at the jobsite for the field worker. Making video capture for VRA as efficient as possible should prevent field workers from avoiding the task.

[0085] Based on the type of deficiency identified by the method 200, the controlling circuitry 114 of the mobile device 102 may be configured to adjust a property of the mobile device 102 to correct the deficiency. The adjustment may be automatic based on the method 200 identifying a deficiency. For example, the method 200 may determine that the video being captured is too dark and that a flash would improve the video. The first prompt, in this case, may be a message indicating that the video is too dark and the controlling circuitry may turn on a flash of the mobile device 102 (for example, where the device 102 is a phone, tablet or digital camera). The prompts 402 may not necessarily indicate that the field worker needs to make an adjustment, but merely provide a warning to the field worker that the mobile device 102 will adjust for the sake of a good quality video. Otherwise, without a prompt 402, the field worker may turn the flash off in this example, not knowing that the action was taken based on the findings of the analysis method 200.

[0086] Video analysis to identify any deficiencies in the video may be performed by an artificial intelligence (Al) model trained to identify deficiencies in a video. A mobile application may be run on the mobile device 102 and capturing a video may be enabled as part of the mobile application, using video capturing software of the mobile device 102. The controlling circuitry 114 of the mobile device 102 may instruct the processor 108 to cause the Al model to perform analysis of the video, with the video captured by the field worker 100 as input.

[0087] Alternatively, the mobile device 102 may issue a request to the server 302 to cause the Al model to perform video analysis. The server 302 may comprise a processor and a memory andmay receive video data from the video being captured by the field worker 100 - from the transceiver 116 of the mobile device 102 for example. The processor 308 at the mobile device 102 may be configured to process the video data from the video being captured before being sent to the server 302, such that the video data in a suitable format for the server 302 to deliver to the Al model as input, or for the server 302 to process for inputting.

[0088] The Al model may be configured to analyse image content of frames of a video. The Al model may be configured to analyse sound content. The Al model may be configured to analyse both image and sound content together - in this way, the Al model may be configured to draw conclusions based on how both the image and sound content in the video interact / relate.

[0089] The Al model is pre-trained to perform video analysis. For some of the parameters the Al model may be trained with labelled data of historical risk assessments, and for some parameters heuristic based analysis, e.g. accelerometer data, may be used. For example, historical risk assessments may be labelled as comprising presence of voice, blurriness, stillness, sufficient duration, etc.

[0090] Figure 5 shows an example data flow for video analysis 500. The video is captured by the video module 104. The video may also include audio, for example captured by the microphone 110. Video data - comprising image data and maybe also comprising audio data - is stored in the memory 112. Video analysis is performed at the processor 108. Alternatively, or additionally, video analysis is performed at the server 302.

[0091] The presence of a deficiency may be identified based on a heuristic marker in the video. A heuristic marker may be a "rule of thumb" as opposed to a direct measurement. For example, where the video being tilted can be measured directly using the accelerometer 604, a heuristic marker may be readings of accelerometer, or change in pixels between the samples of video.

[0092] The output of the video analysis Al model is delivered to a prompt engine 502. The prompt engine 502 is configured to generate a first prompt to be delivered to the field worker 100 at the user interface 400 (on the display 106) and / or via the speaker 110a where the prompt is an audible prompt. A prompt 402 may comprise one or more prompt elements. The prompt engine 502 is configured to draw one or more prompt elements from a prompt element store (which may be at the memory 112) that apply to the received output of the video analysis. A prompt element may be an indicator that a deficiency is present in the video. A prompt element may be a solution to remove a deficiency. Prompt elements may be combined to form a prompt 402. For example, where the output of the video analysis indicates that there is no sound in the video for more than a threshold amount of silence (for example to allow for breathing or thinking or moving to a new sentence, for example) the prompt engine 502 may call upon a prompt element that indicates that there is no sound in the video and may additionally call upon a prompt element indicating that the field worker 100 should narrate the video and / or speak up and / or uncover the microphone 110. The prompt engine 502 may generate a prompt 402 including each relevant prompt element. If the analysis has only been performed once, the prompt 502 may be a first prompt according to embodiments.

[0093] If the analysis has been performed two or more times, the prompt engine 502 may be configured to generate a follow up prompt to the first prompt, by drawing on related prompt elements. The prompt element store may be organised into categories, for example may be organised based on the type of deficiency. In an example, the prompt engine 502 may follow aprompt indicating that the video has no sound with a prompt indicating that the video now has clear sound or that the video still has no sound or that the field worker 100 is doing a good job but should try to speak louder, for example.

[0094] The prompt engine 502 may be configured to generate multiple prompts 402 based on one video analysis output. For example, the video may contain multiple deficiencies. The prompt engine 502 may be configured to rank the deficiencies. The ranking may be based on the order in which the deficiencies occur within the video. The ranking may be based on the severity of the deficiencies. The ranking may be based on the ease with which the field worker 100 can fix the deficiencies, for example. In an example, the output of the video analysis may indicate that the field worker 100 is speaking too quickly and that the video is too dark. The prompt engine 502 may generate a prompt 402 indicating that the video is too dark and / or requesting that the field worker 100 uses a flash. The prompt engine 502 may also generate a prompt 402 indicating that the narration is not clear and / or requesting that the field worker 100 speaks more clearly or slowly. The prompt engine 502 may determine that switching on the flash is a solution that the field worker 100 can apply by making one adjustment. The prompt engine 502 may determine that speaking more slowly or clearly will take effect over time during the video and that it is a physical behaviour of the field worker 100 that needs to be changed. Following the determination by the prompt engine 502 that turning on the flash will be fixed more quickly than changing how the field worker 100 is performing their narration, the prompt regarding the flash may be ranked higher to be presented to the field worker 100. Once the dark video deficiency has been addressed, the field worker 100 may then focus on improving their narration.

[0095] The prompt engine 502 may be configured to prioritise a deficiency that will take longer to fix. Following the above example, the prompt engine 502 may determine that it will take longer to improve narration than it will to activate the flash of the mobile device 102 and so may prioritise a prompt 402 indicating that the narration needs to be improved. This way, the narration may be improved before capturing the video has ended, if improvement is started as soon as possible.

[0096] How the prompt engine 502 prioritises prompts may be based on a set of instructions. For example, mapping of prompts to deficiencies may be used. Ranking may be driven by the order of appearance and or confidence levels of the deficiency being present.

[0097] The Al model may be configured to rank the deficiencies identified in the video and deliver an output to the prompt engine 502 indicating the order in which prompts 402 referring to those deficiencies should be delivered to the field worker 100.

[0098] The prompt engine 502 may be configured to generate a single prompt 402 identifying more than one deficiency. Based on the above example, the prompt engine 502 may generate a prompt 402 using a prompt element indicating that the video is too dark or indicating a solution to a dark video and a prompt element indicating that the narration is too slow or lacking clarity. Presenting a single prompt 402 to the field worker 100 including different prompt elements may increase the efficiency of the field worker 100 in dealing with the prompt 402, versus working their way through a list of individual prompts 402.

[0099] Ideally, the field worker 100 will not be overwhelmed with prompts 402 - for example obscuring a large part of the video field of view and negatively affecting their video capturing (for example creating more deficiencies). Therefore, ranking the prompts 402 may allow theprompt engine 502 to cause only a top-ranking prompt, or a top three prompts, for example to appear on the display 106. The prompts 402 may appear as a list that can be scrolled through or otherwise dynamically interacted with to prevent the prompts 402 from taking over the display 106. As previously noted, a prompt 402 may be interactive in that the field worker 100 may be invited to acknowledge the prompt 402, for example close the prompt 402 using a button on the user interface 400 or a voice command via the microphone 110.[000100]Where only one deficiency is identified in the video, a single prompt 402 may be generated and the ranking of that prompt 402 may - by default - be top. Where the analysis output indicates that no deficiencies are present in the video, the prompt 402 may be a single positive prompt such as "good job". The prompt engine 502 may be configured to generate multiple encouraging prompts 402 to be issued to the field worker 100 at predetermined intervals to encourage the field worker 100 throughout capturing the video.[000101]A prompt 402 may have a predetermined lifetime - for example, being displayed on the display 106 for 1 second or 2 seconds or other times before being automatically removed from the display 106. This may prevent the field worker 100 from getting distracted by a prompt 402. The lifetime of a prompt 402 may be based on the size of the prompt 402, for example based on how long the field worker 100 will need to read a written prompt. The lifetime may be determined by the prompt engine 502. The lifetime may be based on the number of words in the prompt 402. The lifetime may be based on the number of prompt elements making up the prompt 402. Where the prompt 402 is audible or comprises an audible prompt element (some prompts 402 may be both visual and audible, for example the words of a prompt 402 may be spoken from the speaker 110a as well as readable on the display 106) the lifetime may be based on the length of time it takes to issue the audible prompt once, or may comprise repeating the audible prompt an appropriate number of times, for example where the audible prompt comprises a bell or alarm sound or other alert sound.[000102] The prompt engine 502 may be enabled to recall a prompt 402. The prompt engine 502 may be enabled to add or remove a prompt element from a recalled prompt 402 and re-issue the prompt 402 for display. The prompt engine 502 may be configured to re-issue a prompt 402 if the video analysis identifies the same deficiency at a later time. The prompt 402 may be upgraded to a higher position in the ranking of prompts if it is a repeated prompt 402. The prompt engine 502 may add a prompt element to the original prompt, when a prompt is to be repeated, indicating that the prompt is repeated / that a deficiency is persisting / that action by the field worker 100 is urgent.[000103]Video analysis may be ongoing while the video is being captured. Because of this, the priority of prompts 402 may change. If the prompt engine 502 changes the ranking of prompts 402 already appearing to the field worker 100 on the display 106, the ranking change may be represented to the field worker 100 in a noticeable or eye-catching manner, to indicate to the field worker 100 that a prompt 402 has become more important. For example, different colours or numbering may be used to indicate an order change or change of severity. Prompts 402 may be shown to move up or down a list of prompts 402. Prompts 402 may flash or grow in size or otherwise be augmented to catch attention.[000104]0f course, following an encouraging prompt 402 the field worker 100 may change how they are capturing the video and void the encouraging prompt 402. In other words, a deficiency may appear in the video, even when an encouraging prompt 402 is currently shown on the display 106. Then, a prompt 402 indicating the deficiency will take precedence over theencouraging prompt 402. Prompts 402 that indicate a lack of deficiencies - so-called encouraging prompts 402 - may have a predetermined lifetime, so as not to take up space that may shortly be required for a prompt 402 indicating a deficiency, if the video analysis output changes. The predetermined lifetime of an encouraging prompt 402 may be shorter than prompts 402 indicating deficiencies. Encouraging prompts 402 may be automatically recalled if the video analysis output indicates that a deficiency is present - even before a prompt 402 based on the deficiency has been generated.[000105] The field worker 100 may realise that they have caused a deficiency in the video or otherwise identify a deficiency before the video analysis provides an output indicating the deficiency. The field worker 100 may fix the deficiency themselves or it may otherwise be removed. As such, the prompt engine 502 may have generated a prompt 402 that is no longer relevant. To prevent issuing a confusing prompt 402, prompts 402 may have a lead time between being generated and being shown at the display 106. There may be a lead time between video analysis output arriving at the prompt engine 502 and prompt generation. A prompt may be aborted during the lead time if the video analysis determines that the deficiency is no longer present. Likewise, an encouraging prompt 402 may have a lead time and if a deficiency is identified before the end of the lead time, the encouraging prompt 402 is not issued. The lead time may be 0.5 seconds or 1 second or another brief time for example.[000106] Figure 6 shows an example mobile device 102. As well as video information captured by the video module 104 and sound captured by the microphone 110, the mobile device 102 may comprise other modules capable of obtaining data about the surroundings of the mobile device 102, the environment, or how the mobile device 102 is being used for example. The modules may deliver information to the memory 112 to be included as input to the Al model, for example. Any or each of the modules may deliver data to the processor 108 to be processed before storage or before being input into the Al model for example. The mobile device 102 may comprise an accelerometer 604. The accelerometer 604 may be configured to measure motion of the mobile device 102. For example, based on data from the accelerometer 102, it may be determined that the field worker 100 is moving the mobile device 102 quickly or slowly or that they may be tilting the mobile device 102 such that the video captured by video module 104 will be tilted.[000107] In an example, it may be desirable for the video to be taken in portrait (as opposed to landscape or at another angle). The Al model may receive an input indicating that the mobile device 102 is tilted such that the video is being captured in landscape. This may be identified as a deficiency and the prompt engine 502 may be configured to generate a prompt 402 indicating that the mobile device 102 is tilted and / or that the field worker 100 needs to orient the mobile device 102 in a portrait mode.[000108] Based on data from the accelerometer 604, it may be determined that the field worker 100 is moving the mobile device 102 quickly through the jobsite 120 or quickly side to side, for example. This may cause a deficiency in the video, for example blurring or changes to the volume of the field worker's narration (if they are moving the microphone 110 around). Data from the accelerometer 604 may be processed by the processor 108 and provided as an input to the Al model to indicate how the mobile device 102 is being moved around.[000109] The mobile device 102 may comprise a location module 606, which may comprise a positioning receiver such as a global positioning system (GPS) receiver or positioning receiver configured to triangulate the location of the mobile device 102 based on nearby cell towers orknown Wi-Fi locations. Smartphones typically include a location module, for example. Location data taken from the location module 606 may be used to associate the video with a particular location automatically, for example. This may be used to provide known hazard data or expected hazard data to the Al model, for comparison with the video data, to determine if the field worker 100 has missed a known or expected hazard based on the location of the jobsite 120. The location module 606 may also provide confirmation that the mobile device 102 is being moved around at the job site 120. If location data is provided to the Al model as an input and the location of the mobile device 102 does not change during an analysis window 706, or from one analysis window 706 to another, for example, then a deficiency may be identified because the mobile device 102 is stationary.[000110]The mobile device 102 may comprise a lighting module 608 which may comprise a flash, configured to provide light for video. The processor 108 may be configured to determine whether the lighting module 608 has enabled the flash or not during video capture. The status (on or off) of the flash may be delivered to the Al model as an input. If the Al model determines that the video is too dark and that the flash is off, an output to the prompt engine 502 may be to generate a prompt 402 requesting the field worker 100 to turn on the flash. As previously noted, the controlling circuitry 114 may cause the flash to be turned on automatically based on an output from the Al model that the video is too dark. If the Al model has been informed that the flash is on and it determines that the video is too dark, an output to the prompt engine 502 may indicate that a prompt 402 mentioning the flash is not suitable for the situation. The prompt engine 502 may be configured to indicate that the video is too dark, without a solution relating to the flash. Based on an indication that the flash is on, but determining that the video is too dark, the Al model may be configured to output that the camera lens may be covered. The prompt engine 502 may be configured to generate a prompt 402 indicating that the camera lens is covered and / or requesting that the field worker 100 uncovers the lens, for example.[000111]The processor 108 may be configured to process data from any of the modules 604, 606, 608 and 104 for transmission to the server 302. For example, where the Al model is run at the server 302, the server 302 may be configured to request data from any of the modules 604, 606, 608, 104 from the mobile device 102 for input into the model.[000112] Figure 7 shows a video 700, represented by a sequence of frames 702 and audio 704. The audio 704 has been captured as the visual elements of the video have been captured so the audio 704 matches the sequence of frames 702 in time. For the model to analyse whether any deficiencies are present in the video - which will go on to be used for VRA video risk assessment - analysis of a window comprising sufficient information takes place. An analysis window 706 is shown in Figure 7, which comprises a portion of audio 704 and frames 702. The size of the analysis window 706 in Figure 7 is an example only. Typically, in a digital video there is a frame rate of 30 frames per second; however, for the sake of visual representation Figure 7 does not include so many frames 702.[000113] In the example of Figure 7, frames of video 702 that have been captured successfully show a tree. Within the analysis window 706, two frames clearly show a tree, and two frames are completely black. Therefore, there is a deficiency in the video within the analysis window 706. The Al model may determine that the video is too dark or missing completely, and may cause the prompt engine 502 to generate a prompt 402 to that effect. In this example, the video returns (showing a tree) within the analysis window 706, so the Al model may determine that the video was temporarily obscured during the analysis window 706 and that the cause of thedark video is therefore not simply low light and that the camera lens was covered. Therefore, the output of the Al model may cause the prompt engine 502 to generate a prompt indicating that the video is too dark and / or including a solution: "uncover the camera lens" or the like.[000114] The Al model may be configured to identify whether the contents of the frames 702 changes between analysis windows 706. If not, a deficiency may be identified in that the field worker 100 is not moving the camera around, for example.[000115] The size of the analysis window 706 may be determined by the model. The model may be configured to analyse the contents of the or each frame 702 within the analysis window 706 and / or the contents of the audio within the window 706. The model may be configured to analyse two or more analysis windows 706 from the same video. The analysis windows 706 may follow one after another or may overlap (i.e. may include some of the same frames and / or audio). The analysis window 706 may be a sliding window, in that analysis of the window 706 may take place between the start and end of the window 706, but the start and end of the window 706 may not remain fixed. As the video is being captured, the analysis window 706 may progress along the frames 702 and audio 704, for example. The analysis window 706 may have a fixed duration - for example 30 seconds of audio or a total of 30 seconds worth of frames or a fixed number of frames.[000116] The analysis window 706 may have a duration based on the deficiency being identified. There may be multiple analysis windows 706 running in parallel during video capture, for different deficiencies. For example, identifying whether the video is tilted may be a quick determination - for example, performing image processing on the video alongside receiving data from the accelerometer 604. The analysis window 702 may have a duration of 2 seconds, for example. The analysis window 702 may have a duration based on a processing loop length - between receiving the video input to be analysed and outputting an indication that a deficiency is present. Identifying that the video is tilted has a comparatively shorter processing loop length than identifying a problem with the audio, for example, because a larger sample of the audio is needed before a deficiency can be identified.[000117] The duration of an analysis window 706 for identifying a deficiency in the audio may be 20 seconds or 30 seconds, for example. This is because the field worker 100 may be narrating the video suitably at a first time and then take a break from narration that is too long for the VRA, and so it is only once the break in audio reaches a threshold duration that it is determined to be a deficiency. In this way, breathes between sentences and time for the field worker 100 to think about what to say can be accounted for without being identified as deficiencies. It may otherwise put a field worker 100 off capturing video for VRA if prompts 402 are issued too soon or too aggressively and they feel rushed or pressured.[000118] In the example of Figure 7, an audio signal is being received at the start of the analysis window 706 but there is no audio signal in the middle of the window 706 and at the end of the window 706. The audio signal is first missing alongside the video being blank. The audio signal is secondly missing when the visual content is not blank. The model may identify two blank frames of video 702 and two portions of missing / silent audio 704. If the analysis window 706 ended without a visual or audio signal being present, the model may be configured to identify that both the audio and visual signals are missing at the same time, and an output of the model may indicate that the mobile device 102 may be experiencing a fault or that the field worker 100 is covering the camera lens and microphone 110 (for example it may be in their pocket). The prompt engine 502 may be caused to generate a prompt 402 indicating that thevideo cannot be used for VRA and that the field worker 100 needs to try again. In this case, the visual signal returns at the end of the analysis window 706.[000119]Based on the analysis window 706 of Figure 7, the audio signal includes periods of silence. Clearly, the microphone 110 is functioning because audio has been received during the analysis window 706. Therefore, the Al model will determine that the field worker 100 is leaving gaps in their narration. Whether a prompt 402 is generated will be based on whether the gap in audio signal exceeds a threshold time (to allow the field worker 100 to think about what they want to say, for example). The Al model may identify that the analysis window 706 comprises some audio and that may be sufficient to avoid a prompt 402 being generated based on an audio deficiency.[000120] Figure 7 also shows a transcription window 708. Time may be allowed after the analysis window 706 for transcription of audio from within the analysis window 706, such that the model can analyse the transcription. Where there is no or negligible time between the end of a first analysis window 706 and the start of a second analysis window 706, the model may be configured to perform transcription during a transcription window 708 that overlaps in time / runs parallel with the second analysis window 706. Where the analysis window 706 is a sliding window, a sliding transcription window 708 may also be implemented.[000121] A transcription module 610 may be provided, which may be part of the mobile device 102 or may be provided at the server 302. The transcription module 610 may receive audio from the video for transcription. Within the transcription window 708, an audio sample, for example from the analysis window 706, may be provided to the transcription module 610. The transcription module 610 may be configured to output transcription data to the Al model and the Al model may be configured to perform an analysis of (i) word quality and (ii) word count. Transcribing the audio may provide insight into deficiencies in the field worker's 100 narration. For example, although an audio signal may be present in the analysis window 706, it may transpire from the transcription that the words are incomprehensible. The field worker 100 may meet the requirements in that there are no long silences (above a threshold length of time for a silence) but they may be speaking too quickly such that they cannot be understood. For example, the word count may be too high for the duration of the analysis window 706 and the prompt engine 502 may be caused to generate a prompt 402 asking the field worker 100 to speak more slowly. The word count may be too low compared with an expected word count for the duration of the analysis window 706 and the prompt engine 502 may be caused to generate a prompt 402 indicating that the field worker 100 needs to increase their narration. Where the word quality is low, the prompt engine 502 may be configured to generate a prompt 402 asking the field worker 100 to speak more clearly, move closer to the microphone 100, uncover the microphone 110 (for example, the words may be muffled and so of pure quality) or to remove any background noise such as a radio playing.[000122] In an example, an audio signal may be present but only once the transcription module 610 has transcribed the audio does it become clear that the audio does not contain words, but rather noise such as drilling, construction noise or other typical jobsite sounds. The prompt engine 502 may be configured to generate a prompt 402 indicating that the field worker's 100 voice cannot be heard versus background noise.[000123]The transcription of the audio may have a confidence score. The transcription confidence score may be provided as an input into the Al model as the Al model may be configured todetermine whether a particular hazard has been identified and the confidence score may contribute to whether the Al model determines that a hazard has been mentioned or not.[000124] Part of the analysis performed by the Al model may comprise identifying whether one or more expected hazards have been identified by the field worker 100, as a measure of whether they have captured a good video for VRA. For example, if the field worker 100 should have identified a pothole because the jobsite 120 is known to include a pothole, but has failed to capture the pothole in the video, then this indicates that the video is not a complete video of the jobsite's hazards. This may indicate that the field worker 100 is in the wrong place, or has not properly completed the video. An input into the Al model may be a list of known hazards for the jobsite 120. Known hazard data may be requested from the memory 112 or remote storage, for example from the server 302. The Al model may analyse the video to determine whether all of the known hazards have been identified by the field worker 100, either in the visual data or in the audio - for example the field worker 100 may state that they have identified a hazard but have not captured it in the video visually because they have deemed it unsafe to approach, for example.[000125] Expected hazards may not be limited to those known to be present at the jobsite 120, but may include those predicted to be present, for example based on weather data for the local area (which may be obtained via the mobile device 102 or the server 302) or the time of day, for example. Predicted hazards may be includes as an input into the Al model and the Al model may be configured to cross-reference predicted hazards with the video.[000126] Analysis of the video according to the method 200 may take place while the video is being captured, as analysis windows 706 may not be as long as the video as a whole. The analysis windows 706 may be sliding windows, such that analysis can take place throughout the video being captured. Post-recording analysis may also be performed, or performed as an alternative to analysis during the video. The method 200 would be the same whether the video is currently being captured or whether the video has been recorded and then analysed. In an example, steps (i), (ii) and (iii) of the method 200 are performed while the video is being captured. Step (iv) is performed post-recording and the second prompt 402 is generated based on step (iv) post recording. In another example, all of the steps of the method 200 are performed on a pre-recorded video. The first prompt may indicate a first deficiency with the video to the field worker 100 or that the video was of good quality. The second prompt may indicate that the deficiency remained and so that the video needs to be recorded again, or that the deficiency was addressed. If the deficiency was addressed, the video may still need to be recorded again if the deficiency was only addressed towards the end of the video or if other deficiencies remained (which would be conveyed in other prompts, for example). Where the first prompt indicated that the video was of good quality, the second prompt may merely confirm that the video analysis is complete.[000127] Analysis of the video according to the method 200 post-recording may proceed as follows: the input to the Al model may comprise the complete video (all of the sampled frames) and the complete audio, and may comprise a complete transcription of the audio, and may comprise known hazards, and may comprise expected hazards, and may comprise heuristics data. The Al model may be configured to determine if the video is (i) long enough in duration and (ii) clear enough (visually and in the audio). The Al model may be configured to determine if the video is tilted or not. The Al model may be configured to determine if the word count is as expected. The Al model may be configured to determine whether known or expectedhazards are missing from the video, or if a known or expected hazard is missing from the visual data but is mentioned in the narration for example.[000128] Some deficiencies cannot be identified during the video being captured, for example total duration of the video. If the video is too short to use for VRA, this will not be known based only on an analysis window 706 of a duration shorter than the total video duration. Therefore, the method 200 may comprise performing steps (i)-(iii) during capturing of the video, and performing step (iv) during the video being captured to establish whether the field worker 100 responds to the first prompt or to identify new or recurring deficiencies, and / or performing step (iv) once the video has been completed. If step (iv) is performed after the end of the video, the second prompt may be provided once the video has ended, as a concluding remark to the field worker 100.[000129] Additionally, there is provided a computer-implemented method 800 of providing training to a field worker 100 to improve safety on a jobsite 120. The method 800 may be performed using a processor at a remote server - remote meaning remote from the mobile device 102. The method 800 may be performed using a processor 108 at the mobile device 102.[000130] Figure 8 shows steps of the method 800. The method 800 comprises analysing a video recorded using a video recorder of a mobile device 102 at a location where a job is being performed, using a processor 802. Analysing the video comprises identifying one or more quality markers in the video, to determine whether the video is of sufficiently good quality to be used in VRA or whether the video includes any deficiencies. The mobile device 102 is associated with a field worker 100 involved in the job.[000131] The method 800 comprises identifying whether a deficiency is present in the video that would lead to a compromised video risk assessment 804, based on the one or more quality markers. The analysis may be performed by an Al model as described above, according to steps (i) and (ii) of the method 200 described above. If repeat analysis is performed, then step (iv) may also be performed.[000132] The method 800 comprises generating, at the processor, a risk assessment quality score from the video 806, based on the one or more quality markers. The risk assessment quality score may be a numerical indication of how well the field worker 100 has done in capturing the video. The score may be a score out of 100, out of 10, or the like.[000133] The method 800 comprises providing, to the processor, historical risk assessment data 808 labelled with historical quality markers from at least one earlier video recorded by the same field worker 100. The historical risk assessment data may be stored in remote storage, for example at a remote server or in the cloud, or may be stored in the memory 112 of the mobile device 102. The processor 108 may be communicatively coupled to the memory 112 / storage and configured to request the historical risk assessment data from the storage / memory 112.[000134]The method 800 comprises inputting the risk assessment quality score and historical risk assessment data into a risk assessment model 810 configured to generate an up-to-date risk assessment quality score for the field worker 100 and running the model to create an up-to- date risk assessment quality score 812 for the field worker 100. The quality score may be estimated from components like duration, presence of audio, sufficient camera movement, etc. The estimation of the quality score may be made by a quality score Al model trained to estimate such a quality score. The training of the quality score Al model may be based on labelling of historical risk assessment videos.[000135] Further, the method 800 comprises providing an indication of safety performance, based on the up-to-date risk assessment quality score, and enabling training content to be transmitted 814 to the mobile device 102 and / or a user in a supervisory role based on the indication of safety performance. The indication of safety performance may be a numerical score, or may be an assessment of whether the field worker 100 has improved their score since an earlier assessment (based on the historical risk assessment data), or may attribute a grade or status to the field worker 100, for example in connection with a qualification or certification or other grading system. The field worker 100 may be indicated to be "safe", for example, based on the up-to-date risk assessment quality score.[000136] Training content may be provided to the field worker 100 via the user interface 400 of the mobile device 102. Training content may be provided by way of a prompt 402, for example a prompt 402 may include training information such as a tip, in writing, about how to capture a video. A prompt 402 may comprise an external link to a webpage, for example enabling the field worker 100 to view a training video in a web browser on the mobile device 102. Local training content may be available within a user application being used on the mobile device 102, such as a mobile application in which the prompts 402 are being presented. What training content to present to the field worker 100 may be determined based on the identified deficiencies and / or their up-to-date risk assessment quality score. For example, the training may be provided based on a set of if then rules, based on how poor was the performance on certain aspects, e.g. the field worker did not talk at all during VRA, is it a consistent behavior, etc. What training to provide may also depend on the quality score overall. Further, if the quality score Al model estimates that the field worker 100 has not narrated the video properly (for example leaving long pauses of silence or speaking too quickly or the like), an output of the quality score Al model may indicate that the field worker 100 should be provided training on how to narrate a video. The prompt engine 502 may generate a prompt 402 including a link to this relevant training, for example. Based on the output of the Al model, the control circuitry 114 may automatically draw upon relevant training content from a library of training content stored at the mobile device 102, in the memory 112, and cause the training content to appear at the user interface 400 for the field worker 100 to consume, for example.[000137] The method 800 may comprise providing a first visual or audible prompt 402 at the mobile device 102 indicative of the deficiency and / or indicative of a solution to remove the deficiency, or providing a first visual or audible prompt 402 at the mobile device 102 indicating that the video is currently free of deficiencies. The method 800 may further comprise analysing the video after the prompt 402 has been provided to identify whether a deficiency is present in the video that would lead to a compromised video risk assessment. In this way, the field worker 100 has been given an opportunity to correct a deficiency (which may or may not be their fault) and so demonstrate that they can adapt their video capturing to provide a good quality video for VRA. The analysis following the prompt 402 being issued can then determine if the field worker 100 has responded to the prompt 402.[000138] The method 800 may comprise providing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed or at a time after the first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies. Of course, if multiple prompts may be issued between the first and second prompts to arrive at a second prompt indicative that the video is currently free of deficiencies. It may take more than one prompt for the field worker 100 to address and remove the deficiency. But, for clarity, the "second" prompt here means a prompt402 issued later than the first prompt 402 and indicating that the video is currently free of deficiencies (irrespective of how many prompts have come between the first and second prompts).[000139]The method 800 may comprise updating the field worker's up-to-date risk assessment quality score based on one or both of the first and second visual or audible prompts. For example, where the first prompt 402 indicates that it is the field worker 100 who is causing the deficiency - for example by failing to narrate the video properly - their up-to-date risk assessment quality score may go down. If the audio improves and a follow-up prompt indicates that the narration is good, then the field worker's 100 score may go up. If the first prompt 402 is encouraging, then the up-to-date risk assessment quality score may go up, or may stay the same (i.e. may not be reduced). If a deficiency is identified that is not the fault of the field worker 100, and they respond to the first prompt by fixing the deficiency, then their up-to-date risk assessment quality score may go up, for example.[000140] The prompt engine 502 may be provided at the server 302. The prompt engine 502 may be configured to deliver a prompt 402 to different mobile devices - not just to the mobile device 102 of the field worker 100 (or associated with the field worker 100 when the video is captured). The prompt engine 502 may be configured to generate a prompt to be delivered to a different device from the mobile device 102 that captured / is capturing the video. For example, following the method 800 determining an up-to-date risk assessment quality score for a particular field worker 100, a prompt may be issued to a computing device of a user in a supervisory role. The prompt 402 may comprise the field worker's up-to-date risk assessment quality score. The prompt 402 may comprise an indication of or a link to relevant training content. This may enable the supervisory user to proceed with training the field worker 100 accordingly.[000141] The supervisory user may be responsible for multiple field workers. As such, the up-to-date risk assessment quality scores of multiple field workers may be of interest to the supervisory user (e.g. a jobsite manager).[000142] The method 800 may comprise at a remote server 302 identifying the up-to-date risk assessment quality score as belonging to a field worker 100 in a defined team of field workers; and at the remote server 302, automatically aggregating the up-to-date risk assessment quality score with at least one risk assessment quality score of another field worker belonging to the same defined team of field workers to create an up-to-date team risk assessment quality score.[000143] The team risk assessment score may be presented to the supervisory user at their computing device, for example at a user interface with which they can interact. The team risk assessment score may be presented based on which deficiencies are most commonly identified across the team when each field worker's video is analysed, such that the most common deficiencies can be reviewed in training. Based on the team risk assessment score, the supervisory user may be able to perform required training with the team to improve everyone's ability to capture a good quality video for VRA.[000144]The embodiments disclosed herein can be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements.

Claims

CLAIMS1. A method of machine-analysed video risk assessment, the method comprising the following steps:(i) analysing (201) a video recorded using a video recorder of a mobile device (102) at a location where a job is being performed;(ii) identifying (202) whether a deficiency is present in the video that would lead to a compromised video risk assessment;(iii) issuing (203a, 203b) a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency;(iv) analysing (204) the video after the prompt has been issued to identify (204a) whether a deficiency is present in the video that would lead to a compromised video risk assessment.

2. The method of to claim 1, further comprising:(v) issuing (205) a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed.

3. The method of claim 2, further comprising: at step (iii) instead of issuing (203a, 203b) a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency, issuing (203) a first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; and at step (v) instead of issuing (205) a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed, issuing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies at a time after the first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies.

4. The method of any one of claims 1-3, wherein steps (i) to (iii) are performed while the video is being recorded.

5. The method of claim 4, wherein step (iv) is also performed while the video is being recorded.

6. The method of any one of claims 1-5, wherein steps (i) and (iv) are performed at the mobile device (102).

7. The method of any one of claims 1-5 wherein steps (i) and (iv) are performed remote from the mobile device (102) at a remote server (302).

8. The method of claim 7, wherein the remote server (302) is a cloud server.

9. The method of any preceding claim, wherein a field of view of the video being recorded is displayed on a screen (106) of the mobile device (102) for a user to view while the video is being recorded and wherein the visual or audible prompt is a message displayed on the screen10. The method of any preceding claim, wherein at least one of steps (i), (ii) and (iv) are performed using an artificial intelligence model trained to identify one or more types of deficiency in a risk assessment video.

11. The method of any preceding claim, wherein at least one of steps (i), (ii) and (iv) are performed using at least one heuristic marker to identify the presence or absence of a deficiency.

12. The method of any preceding claim, wherein the visual or audible prompt comprises a quality indicator of the video.

13. The method of claim 12, wherein the quality indicator is a message recommending that the user produces a new video.

14. A video risk assessment system (300) comprising: a video recorder; one or more processors; and one or more memories comprising program code instructions for implementing the method according to any one of claims 1-13, when executed on the one or more processors.

15. A non-transitory computer-readable storage medium having stored thereon instructions for implementing the method according to any one of claims 1-13, when executed on one or more devices having processing capabilities.

16. A computer-implemented method of providing training to a field worker to improve safety on a jobsite, the method comprising: analysing a video recorded using a video recorder of a mobile device at a location where a job is being performed, wherein analysing the video comprises identifying one or more quality markers and wherein the mobile device is associated with a field worker involved in the job; identifying whether a deficiency is present in the video that would lead to a compromised video risk assessment, based on the one or more quality markers; generating a risk assessment quality score from the video, based on the one or more quality markers; providing historical risk assessment data labelled with historical quality markers from at least one earlier video recorded by the same field worker; inputting the risk assessment quality score and historical risk assessment data into a risk assessment model configured to generate an up-to-date risk assessment quality score for the field worker and running the model to create an up-to-date risk assessment quality score for the field worker; providing an indication of safety performance, based on the up-to-date risk assessment quality score, and enabling training content to be transmitted to the mobile device and / or a user in a supervisory role based on the indication of safety performance.

17. The computer-implemented method of claim 16, further comprising:providing a first visual or audible prompt at the mobile device indicative of the deficiency and / or indicative of a solution to remove the deficiency, or providing a first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; analysing the video after the prompt has been provided to identify whether a deficiency is present in the video that would lead to a compromised video risk assessment; providing a second visual or audible prompt at the mobile device indicative that the video is currently free of deficiencies once the deficiency has been removed or at a time after the first visual or audible prompt at the mobile device indicating that the video is currently free of deficiencies; and updating the field worker's up-to-date risk assessment quality score based on one or both of the first and second visual or audible prompts.

18. The computer-implemented method of claim 16 or claim 17, wherein the one or more quality markers are identified using an Al model trained to analyse a video and identify the one or more quality markers in image and / or audio information of the video.

19. The computer-implemented method of any one of claims 16-18, wherein the indication of safety performance is provided to the field worker as a visual or audible prompt at the mobile device.

20. The computer-implemented method of any one of claims 16-19, further comprising: at a remote server, identifying the up-to-date risk assessment quality score as belonging to a field worker in a defined team of field workers; and at the remote server, automatically aggregating the up-to-date risk assessment quality score with at least one risk assessment quality score of another field worker belonging to the same defined team of field workers to create an up-to-date team risk assessment quality score.

21. A non-transitory computer-readable storage medium having stored thereon instructions for implementing the computer-implemented method according to any one of claims 16-20, when executed on one or more devices having processing capabilities. a

Citation Information

Patent Citations

  • Adapting workers safety procedures based on inputs from an automated debriefing system

    US20200202472A1

  • Capturing diagnosable video content using a client device

    US20230144621A1