Safety assessment method and system based on construction area monitoring
By acquiring and fusing semantic features of image and audio data in the construction area, the problem of low reliability in construction safety assessment is solved, and a more reliable construction environment safety assessment is achieved.
Patent Information
- Application Number
- CN202511565285.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-30
- Publication Date
- 2026-02-13
AI Technical Summary
In existing technologies, the reliability of safety assessments of construction areas is not high, mainly because the potential semantic information in the collected data is not fully utilized.
By acquiring image and audio data of the target and adjacent grid cells in the construction area, semantic feature mining and fusion are performed to form fused features, and feature restoration is performed to obtain construction safety assessment data.
This improves the reliability of construction safety assessment data. By integrating image and audio features, it can more accurately reflect the impact of equipment distribution and operation on the construction environment.
Smart Images

Figure CN121527697A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data analysis, in particular to a safety evaluation method and system based on construction area monitoring. BACKGROUND
[0002] The traditional engineering safety monitoring method mainly relies on artificial periodic inspection and written report. Project managers need to spend a lot of time collecting and sorting safety data of each construction area, and the real-time and accuracy of the data are difficult to guarantee. At the same time, due to the large scale of the project and the dispersion of the construction area, it is difficult for the on-site management personnel to effectively cover and monitor each construction point, resulting in that some problems cannot be discovered and handled in time. In addition, the responsibility division of different construction areas is not clear enough, and the responsibility is not clear when problems occur, which affects the overall management efficiency of the project. In the prior art, first, the project construction area can be divided into grids, and each grid is responsible for collecting safety data, second, the collected data is sent to the cloud server for storage and analysis through the data transmission module. In the cloud server, the data analysis module analyzes the safety data respectively, and provides the analysis result to the decision support module, finally, the visual display module presents the safety analysis result to the user in an intuitive form, helping the project managers to understand the project status in time and make reasonable decisions.
[0003] However, the inventors have found that in the prior art, for cloud analysis, the potential semantic information in the collected data is not fully utilized, so that the reliability of the analysis result is relatively low, and therefore, there is a problem of relatively low reliability of safety evaluation. SUMMARY
[0004] Therefore, the purpose of the present application is to provide a safety evaluation method and system based on construction area monitoring to improve the problem of relatively low reliability of safety evaluation in the prior art.
[0005] To achieve the above-mentioned purpose, the present application adopts the following technical solutions: A safety evaluation method based on construction area monitoring, comprising: obtaining a target construction image and a target construction audio of a target grid unit in a construction area, and obtaining a neighboring construction image and a neighboring construction audio of a neighboring grid unit of the target grid unit, wherein the target construction image is used at least to reflect equipment distribution information in the target grid unit, and the target construction audio is used at least to reflect equipment running information in the target grid unit; respectively performing semantic feature mining on the target construction image and the neighboring construction image to form a target construction image feature and a neighboring construction image feature; respectively mining semantic features from the target construction image and the adjacent construction image to form target construction image features and adjacent construction image features; fusing the target construction image features and the target construction audio features to form target construction fusion features, and fusing the adjacent construction image features and the adjacent construction audio features to form adjacent construction fusion features; transferring the adjacent construction fusion features into the target construction fusion features to form transferred construction fusion features; performing feature restoration on the transferred construction fusion features to obtain construction safety assessment data, wherein the construction safety assessment data is used to reflect a construction environment safety degree of the target grid unit.
[0006] In a preferred selection of the present application, in the above-mentioned safety assessment method based on construction area monitoring, the step of respectively mining semantic features from the target construction image and the adjacent construction image to form target construction image features and adjacent construction image features comprises: respectively performing convolution mining on the target construction image and the adjacent construction image to form target image convolution features and adjacent image convolution features; performing first deep mining on the target image convolution features to form target construction image features, wherein in the process of first deep mining, feature focusing is performed on the target image convolution features based on local semantic features in the target construction image; performing second deep mining on the adjacent image convolution features to form adjacent construction image features, wherein in the process of second deep mining, feature focusing is performed on the adjacent image convolution features based on local semantic features in the adjacent construction image.
[0007] In a preferred selection of the present application, in the above-mentioned safety assessment method based on construction area monitoring, the step of performing first deep mining on the target image convolution features to form target construction image features comprises: performing connected domain identification on the target construction image, and for each connected domain in the result of connected domain identification, other regions outside the connected domain in the target construction image are masked to form a corresponding target connected domain image; respectively performing convolution mining on each target connected domain image to form each target connected domain convolution feature; respectively focusing the target image convolution features based on each target connected domain convolution feature to form each target image focusing feature; respectively adjusting the target image convolution features based on a focusing importance distribution represented by each target image focusing feature to form each target image adjusted feature; Summation or mean value calculation is performed on each of the target image adjustment features to form target construction image features.
[0008] In a preferred selection of the present application, in the above-mentioned safety assessment method based on construction area monitoring, the step of performing second deep mining on the adjacent image convolution features to form adjacent construction image features comprises: Connected domain identification is performed on the adjacent construction image, and for each connected domain in the connected domain identification result, other regions outside the connected domain are masked in the adjacent construction image to form a corresponding adjacent connected domain image; Convolution mining is respectively performed on each of the adjacent connected domain images to form each adjacent connected domain convolution feature; Based on each of the adjacent connected domain convolution features, correlation focusing is performed on the adjacent image convolution features to form each adjacent image focusing feature; Based on the focusing importance distribution represented by each of the adjacent image focusing features, feature adjustment is performed on the adjacent image convolution features to form each adjacent image adjustment feature; Summation or mean value calculation is performed on each of the adjacent image adjustment features to form adjacent construction image features.
[0009] In a preferred selection of the present application, in the above-mentioned safety assessment method based on construction area monitoring, the step of performing semantic feature mining on the target construction audio and the adjacent construction audio respectively to form target construction audio features and adjacent construction audio features comprises: Convolution mining is respectively performed on the target construction audio and the adjacent construction audio to form target audio convolution features and adjacent audio convolution features; Third deep mining is performed on the target audio convolution features to form target construction audio features, wherein in the third deep mining, feature focusing is performed on the target audio convolution features based on the spectral semantic features in the target construction audio; Fourth deep mining is performed on the adjacent audio convolution features to form adjacent construction audio features, wherein in the fourth deep mining, feature focusing is performed on the adjacent audio convolution features based on the spectral semantic features in the adjacent construction audio.
[0010] In a preferred selection of the present application, in the above-mentioned safety assessment method based on construction area monitoring, the step of performing third deep mining on the target audio convolution features to form target construction audio features comprises: performing Fourier transform on the target construction audio to form a target audio spectrum graph, and determining a fundamental frequency and each harmonic frequency from the target audio spectrum graph; masking information corresponding to each frequency other than the fundamental frequency in the target audio spectrum graph to form a target masked spectrum graph, and for each harmonic frequency, masking information corresponding to each frequency other than the harmonic frequency in the target audio spectrum graph to form a corresponding target masked spectrum graph; performing convolution mining on each target masked spectrum graph to form a convolution feature of each target spectrum graph; performing correlation focusing on the target audio convolution feature based on each target spectrum graph convolution feature to form a target audio focusing feature; performing feature adjustment on the target audio convolution feature based on a focusing importance distribution represented by each target audio focusing feature to form a target audio adjusted feature; performing summation calculation or mean value calculation on each target audio adjusted feature to form a target construction audio feature.
[0011] In the preferred selection of the present application, in the safety assessment method based on construction area monitoring, the step of performing fourth deep mining on the adjacent audio convolution feature to form an adjacent construction audio feature comprises: performing Fourier transform on the target construction audio to form a target audio spectrum graph, and determining a fundamental frequency and each harmonic frequency from the target audio spectrum graph; masking information corresponding to each frequency other than the fundamental frequency in the target audio spectrum graph to form a target masked spectrum graph, and for each harmonic frequency, masking information corresponding to each frequency other than the harmonic frequency in the target audio spectrum graph to form a corresponding target masked spectrum graph; performing convolution mining on each target masked spectrum graph to form a convolution feature of each target spectrum graph; performing correlation focusing on the target audio convolution feature based on each target spectrum graph convolution feature to form a target audio focusing feature; performing feature adjustment on the target audio convolution feature based on a focusing importance distribution represented by each target audio focusing feature to form a target audio adjusted feature; performing summation calculation or mean value calculation on each target audio adjusted feature to form a target construction audio feature.
[0012] In a preferred selection of the present application, in the safety evaluation method based on construction area monitoring, the step of fusing the target construction image features and the target construction audio features to form target construction fusion features, and fusing the adjacent construction image features and the adjacent construction audio features to form adjacent construction fusion features, comprises: fusing the target construction image features into the target construction audio features based on an attention mechanism to form target image attention features, and fusing the target construction audio features into the target construction image features based on an attention mechanism to form target audio attention features, and performing sum calculation or mean calculation on the target image attention features and the target audio attention features to form target construction fusion features; fusing the adjacent construction image features into the adjacent construction audio features based on an attention mechanism to form adjacent image attention features, and fusing the adjacent construction audio features into the adjacent construction image features based on an attention mechanism to form adjacent audio attention features, and performing sum calculation or mean calculation on the adjacent image attention features and the adjacent audio attention features to form adjacent construction fusion features.
[0013] In a preferred selection of the present application, in the safety evaluation method based on construction area monitoring, the step of fusing the target construction image features and the target construction audio features to form target construction fusion features, and fusing the adjacent construction image features and the adjacent construction audio features to form adjacent construction fusion features, comprises: mapping the target construction fusion features to form a first focus importance distribution, and mapping the adjacent construction fusion features to form a second focus importance distribution; performing sum calculation or mean calculation on the first focus importance distribution and the second focus importance distribution to form a target focus importance distribution; based on the target focus importance distribution, adjusting the target construction fusion features to form a transmission construction fusion feature.
[0014] Based on the above, the present application further provides a safety evaluation system based on construction area monitoring, comprising: a memory for storing a computer program; a processor connected with the memory, for executing the computer program stored in the memory to realize the safety evaluation method based on construction area monitoring.
[0015] The safety assessment method and system based on construction area monitoring provided in this application first acquires the target construction image and target construction audio of the target grid cell in the construction area, and acquires the neighboring construction images and neighboring construction audio of the neighboring grid cells of the target grid cell; secondly, semantic feature mining is performed on the target construction image and the neighboring construction image respectively to form target construction image features and neighboring construction image features; then, semantic feature mining is performed on the target construction audio and the neighboring construction audio respectively to form target construction audio features and neighboring construction audio features; then, the target construction image features and the target construction audio features are fused to form target construction fused features, and the neighboring construction image features and the neighboring construction audio features are fused to form neighboring construction fused features; further, the neighboring construction fused features are transferred to the target construction fused features to form transferred construction fused features; finally, feature restoration is performed on the transferred construction fused features to obtain construction safety assessment data. Based on the above, on the one hand, by fusing image and audio features, the latent semantic information of both image and audio dimensions can be integrated, resulting in construction fusion features that can characterize the latent semantic information of equipment distribution and equipment operation. Since equipment distribution and operation directly affect the safety of the construction environment, the reliability of the construction safety assessment data obtained from feature reconstruction based on the construction fusion features can be guaranteed. On the other hand, since not only the latent semantic information of the target grid cell is mined, but also that of neighboring grid cells, the reliability of the obtained construction safety assessment data can be further improved, especially since the construction environment of neighboring grid cells directly affects the construction environment of the target grid cell. Therefore, the safety assessment scheme provided in this application can improve the problem of relatively low reliability in existing safety assessments. Attached Figure Description
[0016] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings.
[0017] Figure 1 This is a structural block diagram of a safety assessment system based on construction area monitoring provided in an embodiment of this application.
[0018] Figure 2 This is a flowchart illustrating the safety assessment method based on construction area monitoring provided in an embodiment of this application.
[0019] Figure 3 This is a schematic diagram of the first depth mining provided for an embodiment of this application.
[0020] Figure 4 This is a schematic diagram of the second depth mining provided in an embodiment of this application.
[0021] Figure 5 A third deep mining diagram provided for the embodiments of the present application.
[0022] Figure 6 A fourth deep mining diagram provided for the embodiments of the present application. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described below in detail with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application.
[0024] As shown in Figure 1 The embodiments of the present application provide a safety evaluation system based on construction area monitoring, which can include a memory and a processor.
[0025] In detail, the memory and the processor are directly or indirectly electrically connected to realize data transmission or interaction. For example, the memory and the processor can be electrically connected through one or more communication buses or signal lines. The processor is used to execute the executable computer program stored in the memory to realize the safety evaluation method based on construction area monitoring provided by the embodiments of the present application.
[0026] Optionally, the memory can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0027] Moreover, the processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.
[0028] It can be understood that,Figure 1 The structure shown is only schematic, and the safety assessment system based on construction area monitoring can also include more or fewer components than those shown, or have a different configuration of components than those shown, such as also including a communication unit for information interaction with other devices (such as image sensors, sound sensors, etc.). Figure 1 Figure 1 The structure shown is only schematic, and the safety assessment system based on construction area monitoring can also include more or fewer components than those shown, or have a different configuration of components than those shown, such as also including a communication unit for information interaction with other devices (such as image sensors, sound sensors, etc.).
[0029] In combination Figure 2 , the embodiments of the present application also provide a safety assessment method based on construction area monitoring, which can be applied to the safety assessment system based on construction area monitoring described above. The method steps defined by the processes related to the safety assessment method based on construction area monitoring can be implemented by the safety assessment system based on construction area monitoring (hereinafter referred to as safety assessment system). The specific processes shown below will be described in detail. Figure 2
[0030] Step S110, obtaining a target construction image and a target construction audio of a target grid unit in a construction area, and obtaining a neighboring construction image and a neighboring construction audio of a neighboring grid unit of the target grid unit.
[0031] In the embodiments of the present application, the safety assessment system can obtain a target construction image and a target construction audio of a target grid unit in a construction area, and obtain a neighboring construction image and a neighboring construction audio of a neighboring grid unit of the target grid unit. The target construction image is at least used to reflect the equipment distribution information in the target grid unit, for example, the target grid unit can be information collected by an image collection device, so as to obtain the target construction image. The target construction audio is at least used to reflect the equipment running information in the target grid unit, for example, the target grid unit can be information collected by a sound sensor, so as to obtain the target construction audio. In addition, the neighboring grid unit can refer to the grid unit adjacent to the target grid unit, or the grid unit with a distance less than a threshold value from the target grid unit, and the threshold value can be selected according to actual needs. Moreover, the specific size of the target grid unit and the neighboring grid unit is not limited, and can also be selected according to actual needs.
[0032] Step S120, respectively performing semantic feature mining on the target construction image and the neighboring construction image to form a target construction image feature and a neighboring construction image feature.
[0033] In this embodiment, after obtaining the target construction image and the adjacent construction images, the safety assessment system can perform semantic feature mining on the target construction image and the adjacent construction images respectively, forming target construction image features and adjacent construction image features. That is, it can mine the latent semantic information in the target construction image to obtain target construction image features, and it can mine the latent semantic information in the adjacent construction images to obtain adjacent construction image features. Furthermore, it should be noted that in this embodiment, the features representing each latent semantic information can be represented as vectors or matrices.
[0034] Step S130: Semantic feature mining is performed on the target construction audio and the adjacent construction audio respectively to form target construction audio features and adjacent construction audio features.
[0035] In this embodiment of the application, after obtaining the target construction audio and the adjacent construction audio, the safety assessment system can perform semantic feature mining on the target construction audio and the adjacent construction audio respectively to form target construction audio features and adjacent construction audio features. That is, it can mine the potential semantic information in the target construction audio to obtain target construction audio features, and it can mine the potential semantic information in the adjacent construction audio to obtain adjacent construction audio features.
[0036] Step S140: Fuse the target construction image features and the target construction audio features to form target construction fusion features, and fuse the neighboring construction image features and the neighboring construction audio features to form neighboring construction fusion features.
[0037] In this embodiment, after obtaining the target construction image features, the target construction audio features, the neighboring construction image features, and the neighboring construction audio features, the safety assessment system can fuse the target construction image features and the target construction audio features to form a target construction fusion feature, and fuse the neighboring construction image features and the neighboring construction audio features to form a neighboring construction fusion feature. It should be noted that since the target construction image features and the target construction audio features both belong to target grid cells and have relatively high correlation, they can be fused first to obtain a target construction fusion feature that can represent the potential semantic information in both dimensions. Similarly, since the neighboring construction image features and the neighboring construction audio features both belong to neighboring grid cells and have relatively high correlation, they can be fused first to obtain a neighboring construction fusion feature that can represent the potential semantic information in both dimensions.
[0038] In step S150, the adjacent construction fusion feature is transmitted into the target construction fusion feature to form a transmission construction fusion feature.
[0039] In the embodiments of the present application, after the target construction fusion feature and the adjacent construction fusion feature are obtained, the safety evaluation system can transmit the adjacent construction fusion feature into the target construction fusion feature to form a transmission construction fusion feature. That is, in the process of potential semantic information fusion, the target construction fusion feature is mainly used, and the adjacent construction fusion feature is used as auxiliary information, that is, the related semantic information of the target grid cell is focused on, and the related semantic information of the adjacent grid cell is used as auxiliary information.
[0040] In step S160, the transmission construction fusion feature is feature restored to obtain construction safety evaluation data.
[0041] In the embodiments of the present application, after the transmission construction fusion feature is obtained, the safety evaluation system can perform feature restoration on the transmission construction fusion feature to obtain construction safety evaluation data. The construction safety evaluation data is used to reflect the safety degree of the construction environment of the target grid cell. For example, the construction safety evaluation data can be a 0-1 or 1-10 value. The larger the value, the higher the safety degree of the construction environment. The smaller the value, the lower the safety degree of the construction environment. For example, if there is an abnormal working device, the safety degree of the construction environment is low.
[0042] Based on the above, on the one hand, the image feature and the audio feature are fused, so that the potential semantic information in the image and audio dimensions can be fused, thereby obtaining a construction fusion feature capable of representing the potential semantic information in the device distribution and device operation dimensions. Therefore, the reliability of the construction safety evaluation data obtained by feature restoration based on the construction fusion feature can be ensured. On the other hand, not only the potential semantic information of the target grid cell is mined, but also the potential semantic information of the adjacent grid cell is mined. Therefore, in the case that the construction environment of the adjacent grid cell directly affects the construction environment of the target grid cell, the reliability of the obtained construction safety evaluation data can be further improved. Therefore, the safety evaluation scheme provided by the present application can improve the problem that the reliability of safety evaluation in the prior art is relatively low.
[0043] In the first aspect, for step S110, the specific way of obtaining the image and audio of the construction area is not limited, and can be selected according to actual needs.
[0044] For example, in an alternative implementation, the image and audio of the construction area collected by the image collection device and the sound sensor in real time can be obtained, so that the target construction image, the target construction audio, the adjacent construction image and the adjacent construction audio capable of representing the current construction environment can be obtained.
[0045] For another example, in another alternative implementation, when the construction safety post-supervision evaluation is needed, the stored historical image and audio of the construction area collected by the image collection device and the sound sensor can also be obtained from the corresponding database, so that the target construction image, the target construction audio, the adjacent construction image and the adjacent construction audio capable of representing the historical construction environment can be obtained.
[0046] The second aspect is that the specific way of performing semantic feature mining on the target construction image and the adjacent construction image respectively in step S120 is not limited, and can be selected according to actual needs.
[0047] For example, in an alternative implementation, in order to mine important semantic information in the process of semantic feature mining, so that the target construction image features and the adjacent construction image features formed have high semantic representation ability, the above step S120 can further include steps S121, S122 and S123, and the specific contents are as follows.
[0048] Step S121: performing convolution mining on the target construction image and the adjacent construction image respectively to form target image convolution features and adjacent image convolution features.
[0049] In the embodiments of the present application, the target construction image and the adjacent construction image can be subjected to convolution mining respectively to form target image convolution features and adjacent image convolution features. For example, the target construction image can be processed by a convolution neural network to obtain target image convolution features, and similarly, the adjacent construction image can also be processed by a convolution neural network to obtain adjacent image convolution features.
[0050] Step S122: performing first deep mining on the target image convolution features to form target construction image features.
[0051] In the embodiment of the present application, after obtaining the target image convolution feature, the target image convolution feature can be subjected to first deep mining to form a target construction image feature. In the first deep mining, the target image convolution feature can be subjected to feature focusing based on the local semantic features in the target construction image. For example, some important semantic information, such as semantic information that can represent the construction environment, can be mined from the target construction image. Then, the target image convolution feature can be guided based on the semantic information in the deep mining process, so that the target construction image feature formed can tend to represent the important semantic information of the construction environment.
[0052] In step S123, the adjacent image convolution feature is subjected to second deep mining to form an adjacent construction image feature.
[0053] In the embodiment of the present application, after obtaining the adjacent image convolution feature, the adjacent image convolution feature can be subjected to second deep mining to form an adjacent construction image feature. In the second deep mining, the adjacent image convolution feature can be subjected to feature focusing based on the local semantic features in the adjacent construction image. For example, some important semantic information, such as semantic information that can represent the construction environment, can be mined from the adjacent construction image. Then, the adjacent image convolution feature can be guided based on the semantic information in the deep mining process, so that the adjacent construction image feature formed can tend to represent the important semantic information of the construction environment.
[0054] It can be understood that the specific manner of first deep mining of the target image convolution feature in step S122 is not limited. For example, in an alternative embodiment, in order to sufficiently mine the semantic information that affects the safety of the environment, step S122 can further include steps S122a, S122b, S122c, S122d and S122e, and the specific contents are as follows (in combination with the description of step S122). Figure 3
[0055] In step S122a, the target construction image is subjected to connected component identification, and for each connected component in the result of connected component identification, other regions outside the connected component in the target construction image are masked to form a corresponding target connected component image.
[0056] In the embodiments of the present application, the target construction image can be subjected to connected domain identification, and for each connected domain in the result of the connected domain identification, other regions in the target construction image except the connected domain are masked to form a corresponding target connected domain image. That is, one target connected domain image can be used to represent one device or one person, or one building pit. In addition, it should be noted that the identification of the connected domain can be implemented by using any existing image algorithm, and is not limited here.
[0057] In step S122b, convolution mining is performed on each target connected domain image to form a target connected domain convolution feature.
[0058] In the embodiments of the present application, after obtaining the target connected domain image, convolution mining can be performed on each target connected domain image to form a target connected domain convolution feature. The convolution mining can be implemented by using a convolution neural network.
[0059] In step S122c, based on each target connected domain convolution feature, the target image convolution feature is subjected to associated focusing to form a target image focusing feature.
[0060] In the embodiments of the present application, after obtaining the target connected domain convolution feature, the target image convolution feature can be subjected to associated focusing based on each target connected domain convolution feature to form a target image focusing feature. That is, since the target connected domain convolution feature is formed by semantic mining of the objects such as devices, persons or building pits that affect the construction environment, the target image focusing feature formed by the associated focusing of the target connected domain convolution feature can focus on representing the semantic information of the objects such as devices, persons or building pits that affect the construction environment. The associated focusing can be implemented by using an attention mechanism, for example, the target image convolution feature can be subjected to cross-attention processing based on the target connected domain convolution feature to form a corresponding target image focusing feature.
[0061] In step S122d, based on the focusing importance distribution represented by each target image focusing feature, the target image convolution feature is subjected to feature adjustment to form a target image adjustment feature.
[0062] In the embodiments of the present application, after the target image focus features are obtained, the target image convolution features can be adjusted based on the focus importance distribution represented by each target image focus feature, to form a target image adjusted feature. That is, since the target image focus features can represent the semantic information of the object affecting the construction environment, the feature parameters at different positions can represent the importance, and based on this, the focus importance distribution representing the importance can be obtained by further mapping the target image focus features, such as using a sigmiod function, and based on the focus importance distribution, the target image convolution features can be adjusted, such as by bit-by-bit multiplication, to obtain a corresponding target image adjusted feature.
[0063] In step S122e, the target image adjusted features are summed or averaged to form a target construction image feature.
[0064] In the embodiments of the present application, after the target image adjusted features are obtained, the target image adjusted features can be summed or averaged to form a target construction image feature, that is, the target image adjusted features are fused.
[0065] It can be understood that the specific manner of first deep mining of the adjacent image convolution features in step S123 is not limited, for example, in an alternative embodiment, in order to fully mine the semantic information affecting the environmental safety, step S123 can further include steps S123a, S123b, S123c, S123d and S123e, and the specific contents are as follows (combined with the description of steps S121a-S121d). Figure 4
[0066] In step S123a, connected component labeling is performed on the adjacent construction image, and for each connected component in the connected component labeling result, other regions outside the connected component are masked in the adjacent construction image to form a corresponding adjacent connected component image.
[0067] In the embodiments of the present application, the adjacent construction image can be subjected to connected component labeling, and for each connected component in the connected component labeling result, other regions outside the connected component are masked in the adjacent construction image to form a corresponding adjacent connected component image. That is, an adjacent connected component image can be used to represent an equipment or a person, or a building pit. In addition, it should be noted that the connected component labeling (CCL) can be implemented by using any existing image algorithm, and is not limited here.
[0068] Step S123b, respectively, on each of the adjacent connected domain image convolution mining, forming each adjacent connected domain convolution features.
[0069] In the embodiment of the application, after obtaining the adjacent connected domain image, each of the adjacent connected domain image can be respectively convolution mining, forming each adjacent connected domain convolution features. Wherein, convolution mining can be realized by convolution neural network.
[0070] Step S123c, respectively, based on each of the adjacent connected domain convolution features, the adjacent image convolution features are associated focus, forming each adjacent image focus features.
[0071] In the embodiment of the application, after obtaining the adjacent connected domain convolution features, each of the adjacent connected domain convolution features can be respectively based on, the adjacent image convolution features are associated focus, forming each adjacent image focus features. That is, because the adjacent connected domain convolution features are formed by semantic mining for the device, the person or the building pit and other objects affecting the construction environment, therefore, through the adjacent connected domain convolution features associated focus, can make the adjacent image focus features formed by focusing on the semantic information of the device, the person or the building pit and other objects affecting the construction environment. Wherein, the association focus can be realized by attention mechanism, for example, can be based on the adjacent connected domain convolution features on the adjacent image convolution features cross attention processing, thereby forming the corresponding adjacent image focus features.
[0072] Step S123d, respectively, based on each of the adjacent image focus features focus importance distribution represented, the adjacent image convolution features are feature adjustment, forming each adjacent image adjustment features.
[0073] In the embodiment of the application, after obtaining the adjacent image focus features, each of the adjacent image focus features can be respectively based on the focus importance distribution represented, the adjacent image convolution features are feature adjustment, forming each adjacent image adjustment features. That is, because the adjacent image focus features can highlight the semantic information of the object affecting the construction environment, therefore, the feature parameters of different positions in it can realize the representation of importance, based on this, the adjacent image focus features can be further mapped, such as through sigmiod function, thereby obtaining the focus importance distribution which can represent the importance, based on this, the focus importance distribution can be based on the adjacent image convolution features feature adjustment, such as bit multiplication, thereby obtaining the corresponding one adjacent image adjustment features.
[0074] Step S123e, summing or averaging each of the adjacent image adjustment features to form an adjacent construction image feature.
[0075] In the embodiments of the present application, after obtaining the adjacent image adjustment features, summing or averaging each of the adjacent image adjustment features to form an adjacent construction image feature, i.e., fusing the adjacent image adjustment features.
[0076] The third aspect needs to be explained for step S130, and the specific way of performing semantic feature mining on the target construction audio and the adjacent construction audio respectively is not limited, which can be selected according to actual needs.
[0077] For example, in an alternative implementation, in order to be able to mine important semantic information in the process of semantic feature mining, so that the target construction audio feature and the adjacent construction audio feature formed have high semantic representation ability, the above-mentioned step S130 can further include steps S131, S132 and S133, and the specific contents are as follows.
[0078] Step S131, respectively performing convolution mining on the target construction audio and the adjacent construction audio to form target audio convolution features and adjacent audio convolution features.
[0079] In the embodiments of the present application, the target construction audio and the adjacent construction audio can be respectively subjected to convolution mining to form target audio convolution features and adjacent audio convolution features. For example, the target construction audio can be processed by a convolutional neural network to obtain target audio convolution features, and similarly, the adjacent construction audio can also be processed by a convolutional neural network to obtain adjacent audio convolution features.
[0080] Step S132, performing third depth mining on the target audio convolution features to form a target construction audio feature.
[0081] In the embodiments of the present application, after obtaining the target audio convolution features, the target audio convolution features can be subjected to third depth mining to form a target construction audio feature. In the process of third depth mining, the target audio convolution features are subjected to feature focusing based on the spectral semantic features in the target construction audio. For example, some important spectral semantic information, such as semantic information that can highlight the representation of the construction environment, can be mined from the target construction audio, and then, based on these spectral semantic information, guidance can be provided in the process of depth mining of the target audio convolution features, so that the target construction audio feature formed can tend to represent the important semantic information of the construction environment.
[0082] Step S133, performing fourth deep mining on the adjacent audio convolution feature to form an adjacent construction audio feature.
[0083] In the embodiment of the present application, after obtaining the adjacent audio convolution feature, fourth deep mining can be performed on the adjacent audio convolution feature to form an adjacent construction audio feature. In the process of the fourth deep mining, the adjacent audio convolution feature is focused based on the spectral semantic feature in the adjacent construction audio. For example, some important spectral semantic information, such as semantic information that can represent the construction environment, can be mined from the adjacent construction audio. Then, based on the spectral semantic information, the process of deep mining of the adjacent audio convolution feature can be guided, so that the adjacent construction audio feature formed can tend to represent the important semantic information of the construction environment.
[0084] It can be understood that the specific manner of performing the third deep mining on the target audio convolution feature in step S132 is not limited. For example, in an alternative embodiment, in order to achieve effective focusing of features in the process of third deep mining, step S132 can further include steps S132a, S132b, S132c, S132d, S132e and S132f, the specific contents of which are described below (in combination with the description of steps S132a-S132f). Figure 5
[0085] Step S132a, performing Fourier transform on the target construction audio to form a target audio spectrum graph, and determining a fundamental frequency and each harmonic frequency from the target audio spectrum graph.
[0086] In the embodiment of the present application, Fourier transform can be performed on the target construction audio to form a target audio spectrum graph, that is, time-domain audio data is converted into a frequency domain, and a fundamental frequency and each harmonic frequency, such as 2 times the fundamental frequency, 3 times the fundamental frequency, etc., are determined from the target audio spectrum graph.
[0087] Step S132b, in the target audio spectrum graph, masking information corresponding to each frequency other than the fundamental frequency to form a target masked spectrum graph, and for each harmonic frequency, in the target audio spectrum graph, masking information corresponding to each frequency other than the harmonic frequency to form a corresponding target masked spectrum graph.
[0088] In the embodiments of the present application, after obtaining the fundamental frequency and each harmonic frequency, the information corresponding to each frequency other than the fundamental frequency in the target audio spectrum diagram can be masked to form a target masked spectrum diagram, i.e., only including information related to the fundamental frequency, and for each harmonic frequency, the information corresponding to each frequency other than the harmonic frequency in the target audio spectrum diagram can be masked to form a corresponding target masked spectrum diagram, i.e., only including information corresponding to the harmonic frequency. It should be noted that different devices usually have specific fundamental frequencies and harmonic modes, and therefore, by mining relevant information, the device and device anomalies can be easily identified.
[0089] In step S132c, each target masked spectrum diagram is respectively subjected to convolution mining to form each target spectrum diagram convolution feature.
[0090] In the embodiments of the present application, after obtaining the target masked spectrum diagram, each target masked spectrum diagram can be respectively subjected to convolution mining to form each target spectrum diagram convolution feature. For example, the target masked spectrum diagram can be processed by a convolutional neural network to obtain a corresponding target spectrum diagram convolution feature.
[0091] In step S132d, based on each target spectrum diagram convolution feature, the target audio convolution feature is respectively subjected to association focusing to form each target audio focusing feature.
[0092] In the embodiments of the present application, after obtaining the target spectrum diagram convolution feature, based on each target spectrum diagram convolution feature, the target audio convolution feature can be respectively subjected to association focusing to form each target audio focusing feature. For example, based on the first target spectrum diagram convolution feature, the target audio convolution feature can be subjected to association focusing to form the first target audio focusing feature. For another example, based on the second target spectrum diagram convolution feature, the target audio convolution feature can be subjected to association focusing to form the second target audio focusing feature. In addition, it should be noted that the association focusing can be realized based on an attention mechanism, as previously described.
[0093] In step S132e, based on the focusing importance distribution represented by each target audio focusing feature, the target audio convolution feature is respectively subjected to feature adjustment to form each target audio adjustment feature.
[0094] In the embodiments of the present application, after obtaining the target audio focusing feature, based on the focusing importance distribution represented by each target audio focusing feature, the target audio convolution feature can be respectively subjected to feature adjustment to form each target audio adjustment feature, as previously described.
[0095] Step S132f, summing or averaging each of the target audio adjustment features to form a target construction audio feature.
[0096] In the embodiments of the present application, after obtaining the target audio adjustment features, summing or averaging each of the target audio adjustment features to form a target construction audio feature, so that the fusion of each of the target audio adjustment features can be realized.
[0097] It can be understood that the specific manner of the third deep mining of the target audio convolution feature in the above step S133 is not limited, for example, in an alternative embodiment, in order to realize effective focusing of the features in the process of the third deep mining, the above step S133 can further include steps S133a, S133b, S133c, S133d, S133e and S133f, and the specific contents are as follows (combined with the description of the above step S133). Figure 6
[0098] Step S133a, performing Fourier transform on the adjacent construction audio to form an adjacent audio spectrum graph, and determining the fundamental frequency and each harmonic frequency from the adjacent audio spectrum graph.
[0099] In the embodiments of the present application, the adjacent construction audio can be subjected to Fourier transform to form an adjacent audio spectrum graph, that is, the time domain audio data is converted into the frequency domain, and the fundamental frequency and each harmonic frequency, such as 2 times the fundamental frequency, 3 times the fundamental frequency, etc., are determined from the adjacent audio spectrum graph.
[0100] Step S133b, in the adjacent audio spectrum graph, masking the information corresponding to each frequency other than the fundamental frequency to form an adjacent masked spectrum graph, and for each harmonic frequency, masking the information corresponding to each frequency other than the harmonic frequency in the adjacent audio spectrum graph to form a corresponding adjacent masked spectrum graph.
[0101] In the embodiments of the present application, after obtaining the fundamental frequency and each harmonic frequency, in the adjacent audio spectrum graph, the information corresponding to each frequency other than the fundamental frequency can be masked to form an adjacent masked spectrum graph, that is, only the information related to the fundamental frequency is included, and for each harmonic frequency, the information corresponding to each frequency other than the harmonic frequency in the adjacent audio spectrum graph is masked to form a corresponding adjacent masked spectrum graph, that is, only the information corresponding to the harmonic frequency is included. It should be noted that different devices usually have specific fundamental frequencies and harmonic modes, and therefore, by mining the relevant information, the identification of the device and the abnormality of the device can be facilitated.
[0102] Step S133c, respectively, each of the adjacent mask spectrum diagram convolution mining, forming each adjacent spectrum diagram convolution features.
[0103] In the embodiment of the application, after obtaining the adjacent mask spectrum diagram, each of the adjacent mask spectrum diagram can be respectively convolution mining, forming each adjacent spectrum diagram convolution features. For example, the adjacent mask spectrum diagram can be processed by a convolutional neural network, thereby obtaining a corresponding adjacent spectrum diagram convolution feature.
[0104] Step S133d, respectively, based on each of the adjacent spectrum diagram convolution features, the adjacent audio convolution features are associated with focusing, forming each adjacent audio focusing features.
[0105] In the embodiment of the application, after obtaining the adjacent spectrum diagram convolution features, each of the adjacent spectrum diagram convolution features can be respectively based on, the adjacent audio convolution features are associated with focusing, forming each adjacent audio focusing features. For example, based on the first adjacent spectrum diagram convolution features, the adjacent audio convolution features are associated with focusing, forming the first adjacent audio focusing features. For another example, based on the second adjacent spectrum diagram convolution features, the adjacent audio convolution features are associated with focusing, forming the second adjacent audio focusing features. In addition, it should be noted that the associated focusing can be realized based on the attention mechanism, as described previously.
[0106] Step S133e, respectively, based on the focusing importance distribution represented by each of the adjacent audio focusing features, the adjacent audio convolution features are feature adjusted, forming each adjacent audio adjustment features.
[0107] In the embodiment of the application, after obtaining the adjacent audio focusing features, each of the adjacent audio focusing features can be respectively based on the focusing importance distribution represented by, the adjacent audio convolution features are feature adjusted, forming each adjacent audio adjustment features, as described previously.
[0108] Step S133f, each of the adjacent audio adjustment features is summed or mean calculated, forming adjacent construction audio features.
[0109] In the embodiment of the application, after obtaining the adjacent audio adjustment features, each of the adjacent audio adjustment features can be summed or mean calculated, forming adjacent construction audio features, so that the fusion of each of the adjacent audio adjustment features can be realized.
[0110] The fourth aspect, for step S140 needs to be explained is that the specific way of fusing image features and audio features is not limited, can be selected according to actual demand.
[0111] For example, in an alternative implementation, in order to make the semantic information of the two dimensions of images and audio fully fused, the above step S140 can further include steps S141 and S142, the specific contents are as follows.
[0112] Step S141, based on the attention mechanism, the target construction image features are fused into the target construction audio features to form target image attention features, and based on the attention mechanism, the target construction audio features are fused into the target construction image features to form target audio attention features, and the target image attention features and the target audio attention features are summed or averaged to form target construction fusion features.
[0113] In the embodiments of the present application, based on the attention mechanism, the target construction image features can be fused into the target construction audio features to form target image attention features, for example, based on the target construction image features, cross-attention processing can be performed on the target construction audio features to obtain target image attention features. And based on the attention mechanism, the target construction audio features are fused into the target construction image features to form target audio attention features, for example, based on the target construction audio features, cross-attention processing can be performed on the target construction image features to obtain target audio attention features. In this way, after bidirectional attention processing, the target image attention features and the target audio attention features can be summed or averaged to form target construction fusion features.
[0114] Step S142, based on the attention mechanism, the adjacent construction image features are fused into the adjacent construction audio features to form adjacent image attention features, and based on the attention mechanism, the adjacent construction audio features are fused into the adjacent construction image features to form adjacent audio attention features, and the adjacent image attention features and the adjacent audio attention features are summed or averaged to form adjacent construction fusion features.
[0115] In the embodiments of the present application, the adjacent construction image features can be fused into the adjacent construction audio features based on an attention mechanism to form adjacent image attention features, for example, the adjacent construction audio features can be processed by cross-attention based on the adjacent construction image features to obtain the adjacent image attention features. In addition, the adjacent construction audio features can be fused into the adjacent construction image features based on an attention mechanism to form adjacent audio attention features, for example, the adjacent construction image features can be processed by cross-attention based on the adjacent construction audio features to obtain the adjacent audio attention features. In this way, after bidirectional attention processing, the adjacent image attention features and the adjacent audio attention features can be calculated by summation or mean value to form the adjacent construction fusion features.
[0116] In the fifth aspect, it needs to be explained that the specific way of transferring the adjacent construction fusion features into the target construction fusion features is not limited and can be selected according to actual needs.
[0117] For example, in an alternative implementation, in order to enable the semantic information in the adjacent construction fusion features to be fused into the target construction fusion features through semantic feature transfer, so as to achieve effective semantic constraint on the target construction fusion features, and further guarantee the semantic representation accuracy of the transferred construction fusion features, the above-mentioned step S150 can further include steps S151, S152 and S153, and the specific contents are as follows.
[0118] Step S151: mapping the target construction fusion features to form a first focus importance distribution, and mapping the adjacent construction fusion features to form a second focus importance distribution.
[0119] In the embodiments of the present application, the target construction fusion features can be mapped to form a first focus importance distribution, and the adjacent construction fusion features can be mapped to form a second focus importance distribution. The mapping can be realized by a sigmiod function, or in other implementations, the target construction fusion features and the adjacent construction fusion features can be linearly mapped respectively to capture the relevant linear semantic relationship, and then the sigmiod function is used to map the results of linear mapping to obtain the first focus importance distribution and the second focus importance distribution respectively. In addition, the linear mapping can be realized by a linear mapping function, such as y=Ax+b, where A is a weight matrix and b is a bias parameter.
[0120] Step S152: calculating the sum or mean value of the first focus importance distribution and the second focus importance distribution to form a target focus importance distribution.
[0121] In the embodiment of the present application, after the first focus importance distribution and the second focus importance distribution are obtained, the first focus importance distribution and the second focus importance distribution can be calculated by summation or mean value to form a target focus importance distribution. That is, in the process of determining importance, not only the semantic information in the adjacent construction fusion feature is considered, but also the semantic information in the target construction fusion feature, that is, not only the data of the adjacent grid cell is relied on, but also the importance represented by the data of the target grid cell itself is combined, so that the reliability of the target focus importance distribution formed can be higher.
[0122] In step S153, the target construction fusion feature is adjusted based on the target focus importance distribution to form a transmission construction fusion feature.
[0123] In the embodiment of the present application, after the target focus importance distribution is obtained, the target construction fusion feature can be adjusted based on the target focus importance distribution, such as multiplied by bit by bit, to realize importance-based weighting, thereby forming a transmission construction fusion feature.
[0124] In the sixth aspect, it needs to be explained that the specific way of feature restoration of the transmission construction fusion feature in step S160 is not limited, and can be selected according to actual needs.
[0125] For example, in an alternative embodiment, step S160 can include the following contents: Firstly, the transmission construction fusion feature can be fully connected to form a corresponding construction fully connected feature, wherein the size of the construction fully connected feature can be 1*1; then, the construction fully connected feature can be identity mapped or activated to obtain construction safety evaluation data.
[0126] In summary, the safety evaluation method and system based on construction area monitoring are provided, first, the target construction image and the target construction audio of the target grid unit in the construction area are obtained, and the adjacent construction image and the adjacent construction audio of the adjacent grid unit of the target grid unit are obtained; secondly, the semantic feature mining is performed on the target construction image and the adjacent construction image respectively to form the target construction image feature and the adjacent construction image feature; then, the semantic feature mining is performed on the target construction audio and the adjacent construction audio respectively to form the target construction audio feature and the adjacent construction audio feature; thereafter, the target construction image feature and the target construction audio feature are fused to form the target construction fusion feature, and the adjacent construction image feature and the adjacent construction audio feature are fused to form the adjacent construction fusion feature; further, the adjacent construction fusion feature is transmitted to the target construction fusion feature to form the transmission construction fusion feature; finally, the transmission construction fusion feature is restored to obtain the construction safety evaluation data. Based on the above, on the one hand, since the image feature and the audio feature are fused, the potential semantic information in the two dimensions of image and audio can be fused, so that the construction fusion feature which can represent the potential semantic information in the two dimensions of equipment distribution and equipment operation is obtained, and thus, since the equipment distribution and the equipment operation will directly affect the safety degree of the construction environment, the reliability of the construction safety evaluation data obtained by restoring the construction fusion feature can be ensured. On the other hand, since the potential semantic information of the target grid unit and the potential semantic information of the adjacent grid unit are mined, in the case that the construction environment of the adjacent grid unit directly affects the construction environment of the target grid unit, the reliability of the obtained construction safety evaluation data can be further improved. Therefore, by using the safety evaluation scheme provided in the present application, the problem of relatively low reliability of safety evaluation in the prior art can be improved.
[0127] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Those skilled in the art can make various modifications and changes to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A safety assessment method based on construction area monitoring, characterized in that, include: Acquire target construction images and target construction audio of target grid cells in the construction area, and acquire neighboring construction images and neighboring construction audio of neighboring grid cells of the target grid cell, wherein the target construction images are used to reflect at least the equipment distribution information in the target grid cell, and the target construction audio is used to reflect at least the equipment operation information in the target grid cell; Semantic feature mining is performed on the target construction image and the neighboring construction images respectively to form target construction image features and neighboring construction image features; Semantic feature mining is performed on the target construction audio and the adjacent construction audio respectively to form target construction audio features and adjacent construction audio features; The target construction image features and the target construction audio features are fused to form a target construction fusion feature, and the neighboring construction image features and the neighboring construction audio features are fused to form a neighboring construction fusion feature; The adjacent construction fusion feature is transferred to the target construction fusion feature to form a transferred construction fusion feature; The construction integration features are restored to obtain construction safety assessment data, which reflects the safety of the construction environment of the target grid cell.
2. The safety assessment method based on construction area monitoring according to claim 1, characterized in that, The step of performing semantic feature mining on the target construction image and the neighboring construction images respectively to form target construction image features and neighboring construction image features includes: Convolution mining is performed on the target construction image and the neighboring construction images respectively to form convolutional features of the target image and convolutional features of the neighboring images; The convolutional features of the target image are subjected to a first depth mining to form target construction image features. During the first depth mining process, feature focusing is performed on the convolutional features of the target image based on the local semantic features in the target construction image. A second depth mining is performed on the convolutional features of the neighboring images to form features of the neighboring construction images. During the second depth mining process, feature focusing is performed on the convolutional features of the neighboring images based on the local semantic features in the neighboring construction images.
3. The safety assessment method based on construction area monitoring according to claim 2, characterized in that, The step of performing a first depth mining on the convolutional features of the target image to form the target construction image features includes: The target construction image is subjected to connected component identification, and for each connected component in the connected component identification result, other areas outside the connected component in the target construction image are masked to form a corresponding target connected component image. Convolution mining is performed on each of the target connected component images to form convolutional features for each target connected component. Based on each of the target connected component convolutional features, the convolutional features of the target image are correlated and focused to form each target image focusing feature; Based on the focus importance distribution represented by the focus features of each target image, feature adjustment is performed on the convolutional features of the target image to form an adjusted feature for each target image. For each of the target image adjustment features, summation or mean calculation is performed to form the target construction image features.
4. The safety assessment method based on construction area monitoring according to claim 2, characterized in that, The step of performing a second depth mining on the convolutional features of the neighboring images to form features of the neighboring construction images includes: Connectivity identification is performed on the adjacent construction images, and for each connected component in the connected component identification result, other areas outside the connected component in the adjacent construction images are masked to form a corresponding adjacent connected component image. Convolution mining is performed on each of the neighboring connected component images to form convolution features for each neighboring connected component. Based on each of the convolutional features of the neighboring connected components, the convolutional features of the neighboring images are correlated and focused to form each neighboring image focusing feature; Based on the focus importance distribution represented by the focus features of each neighboring image, feature adjustment is performed on the convolutional features of the neighboring images to form each neighboring image adjustment feature; Summation or mean calculation is performed on each of the neighboring image adjustment features to form neighboring construction image features.
5. The safety assessment method based on construction area monitoring according to claim 1, characterized in that, The step of performing semantic feature mining on the target construction audio and the adjacent construction audio respectively to form target construction audio features and adjacent construction audio features includes: Convolution mining is performed on the target construction audio and the neighboring construction audio respectively to form target audio convolution features and neighboring audio convolution features; A third depth mining is performed on the target audio convolutional features to form target construction audio features. During the third depth mining process, feature focusing is performed on the target audio convolutional features based on the spectral semantic features in the target construction audio. A fourth depth mining is performed on the convolutional features of the neighboring audio to form neighboring construction audio features. During the fourth depth mining process, feature focusing is performed on the convolutional features of the neighboring audio based on the spectral semantic features in the neighboring construction audio.
6. The safety assessment method based on construction area monitoring according to claim 5, characterized in that, The step of performing a third depth mining on the target audio convolutional features to form target construction audio features includes: The target construction audio is subjected to Fourier transform to form a target audio spectrum diagram, and the fundamental frequency and each harmonic frequency are determined from the target audio spectrum diagram; In the target audio spectrum diagram, the information corresponding to each frequency other than the fundamental frequency is masked to form a target masking spectrum diagram. Also, for each harmonic frequency, the information corresponding to each frequency other than the harmonic frequency is masked in the target audio spectrum diagram to form a corresponding target masking spectrum diagram. Convolution mining is performed on each of the target masking spectrograms to form convolutional features for each target spectrogram. Based on each of the target spectrogram convolutional features, the target audio convolutional features are correlated and focused to form each target audio focusing feature; Based on the focus importance distribution represented by each of the target audio focus features, feature adjustment is performed on the target audio convolution features to form each target audio adjustment feature; For each of the target audio adjustment features, a summation or mean calculation is performed to form the target construction audio feature.
7. The safety assessment method based on construction area monitoring according to claim 5, characterized in that, The step of performing a fourth depth mining on the convolutional features of the neighboring audio to form the neighboring construction audio features includes: The adjacent construction audio is subjected to Fourier transform to form an adjacent audio spectrum diagram, and the fundamental frequency and each harmonic frequency are determined from the adjacent audio spectrum diagram; In the adjacent audio spectrum diagram, the information corresponding to each frequency other than the fundamental frequency is masked to form an adjacent masking spectrum diagram. Also, for each harmonic frequency, the information corresponding to each frequency other than the harmonic frequency is masked in the adjacent audio spectrum diagram to form a corresponding adjacent masking spectrum diagram. Each of the neighboring masking spectrograms is subjected to convolutional mining to form convolutional features for each neighboring spectrogram. Each neighboring audio convolutional feature is associated and focused based on each of the neighboring spectrogram convolutional features to form a neighboring audio focusing feature; Based on the focus importance distribution represented by each of the neighboring audio focus features, feature adjustment is performed on the neighboring audio convolution features to form each neighboring audio adjustment feature; The summation or mean of each of the adjacent audio modulation features is performed to form the adjacent construction audio features.
8. The safety assessment method based on construction area monitoring according to any one of claims 1-7, characterized in that, The steps of fusing the target construction image features and the target construction audio features to form a target construction fusion feature, and fusing the neighboring construction image features and the neighboring construction audio features to form a neighboring construction fusion feature, include: Based on the attention mechanism, the target construction image features are fused into the target construction audio features to form target image attention features, and based on the attention mechanism, the target construction audio features are fused into the target construction image features to form target audio attention features. In addition, the target image attention features and the target audio attention features are summed or averaged to form target construction fusion features. Based on the attention mechanism, the features of the neighboring construction images are fused into the features of the neighboring construction audio to form neighboring image attention features, and based on the attention mechanism, the features of the neighboring construction audio are fused into the features of the neighboring construction images to form neighboring audio attention features. In addition, the neighboring image attention features and the neighboring audio attention features are summed or mean-calculated to form neighboring construction fused features.
9. The safety assessment method based on construction area monitoring according to any one of claims 1-7, characterized in that, The adjacent construction fusion feature is then transferred to the target construction fusion feature. The steps to form the characteristics of construction integration include: The target construction fusion features are mapped to form a first focus importance distribution, and the adjacent construction fusion features are mapped to form a second focus importance distribution; The first focus importance distribution and the second focus importance distribution are summed or mean-calculated to form the target focus importance distribution; Based on the target focus importance distribution, the target construction fusion characteristics are adjusted to form the transfer construction fusion characteristics.
10. A safety assessment system based on construction area monitoring, characterized in that, include: Memory, used to store computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the safety assessment method based on construction area monitoring as described in any one of claims 1-9.