Constricted space quality detection method based on multi-modal edge cloud transmission
Through the multimodal edge cloud transmission system, combined with high-resolution video decoding and efficient data transmission, the problems of low detection efficiency and poor information transmission in confined spaces are solved, and efficient and accurate detection and real-time communication are achieved.
Patent Information
- Application Number
- CN202511324377.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-09-17
AI Technical Summary
In the field of high-end manufacturing, the inspection efficiency of precision components in confined spaces is low, the error rate of manual inspection is high, and the real-time information interaction and high-resolution inspection image transmission are severely delayed, which cannot meet the real-time communication needs.
A multimodal edge cloud transmission system is adopted, combined with a head-mounted camera, a flexible borescope device and a step gap detection device. High-resolution video decoding is achieved through FFmpeg technology, and the 10 Gigabit Ethernet TCP/IP protocol stack is used for data transmission. The CBAM attention mechanism and SIoU loss function are introduced in the YOLOv8 network for micro-defect detection.
It achieves efficient and accurate detection in confined spaces, supports real-time transmission of high-resolution 30FPS video streams, reduces transmission delays, improves the detection rate of micro-defects, and supports data traceability and real-time voice interaction.
Smart Images

Figure CN120833545A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of industrial manufacturing and quality detection, and particularly relates to a quality detection method for a restricted space based on multi-modal edge cloud transmission. BACKGROUND
[0002] In the modern high-end manufacturing field, there are a large number of narrow spaces with limited physical accessibility in the high-density device cluster scene, and the in-situ quality detection of the internal precision components faces technical bottlenecks. The traditional detection mode relies on manual visual inspection and basic measurement tools, which not only is limited by the ergonomics constraints to result in low detection efficiency, but also has a very high misjudgment rate in 0.1mm-level micro-defect identification. Human vision and judgment ability is limited, especially in long-time and high-intensity working conditions, which is more likely to cause fatigue and misjudgment.
[0003] Real-time information interaction in the restricted space detection scene also has significant technical bottlenecks. Due to the physical constraints of the space environment, the two-way communication between the detection personnel and the external control center faces signal attenuation, multipath effect and other interferences, resulting in transmission delay of key detection data, and the real-time return demand of high-resolution detection pictures is particularly prominent. Under the existing technical conditions, it is difficult to meet the dynamic image transmission requirement of 30FPS, resulting in the phenomenon of video stream freezing and mosaic, which seriously affects the remote decision-making efficiency. Therefore, constructing a low-delay and high-reliability communication and transmission system has become a core technical requirement to ensure the safety of the restricted space operation.
[0004] With the progress of science and technology and the development of industrial automation, finding a more efficient, accurate and safe detection method has become a problem to be solved in modern industrial production. This not only requires that the new detection method can adapt to the characteristics of the restricted space, but also needs to have the ability of real-time information transmission and remote collaboration to ensure the smooth progress of the detection work and the accuracy of the detection results. Therefore, the present application proposes an intelligent detection system based on multi-modal edge cloud transmission suitable for restricted space. SUMMARY
[0005] The purpose of the present application is to solve the problems of the prior art and provide a quality detection method for a restricted space based on multi-modal edge cloud transmission, which can solve the problems of low quality detection efficiency, poor information transmission, and inability to communicate and collaborate with external personnel in the restricted space.
[0006] The present application is implemented by the following technical solutions: A quality detection method for a restricted space based on multi-modal edge cloud transmission, comprising the following steps: Step S1: An internal operator carries an edge detection device into a workpiece to be detected, and starts each data acquisition device in the restricted space; Step S2: The wrist display terminal worn by the internal operator displays the video pictures of the head-mounted camera and the hole probe device in real time; Step S3: The video data of the head-mounted camera and the hole probe device inside the restricted space, the step difference gap detection data, and the real-time voice interaction data are transmitted to the external data processing center through TCP / IP; Step S4: The external data processing center performs scratch detection on the received video data, and selectively enters the scratch detection key frame data, video stream data, and step difference gap detection data into the database.
[0007] Preferably, in step S1, the workpiece to be detected is determined, and the internal operator carries each data acquisition device and edge detection device into the narrow space of the workpiece to be detected; each data acquisition device includes a head-mounted camera, a hole probe device, a step difference gap detection device, and a head-mounted earphone; the edge detection device is connected to each data acquisition device as a data transfer station, and the data acquisition devices are used to scan the inside of the workpiece and transmit to the edge detection device; each set of portable device is equipped with a lighting device, and the operator turns on or off the lighting device according to the need to ensure the usability of the collected data.
[0008] Preferably, in step S2, the following steps are included: Step S21: The edge detection device is equipped with a wrist-type small liquid crystal display screen, and the internal operator can view the video pictures of the head-mounted camera and the hole probe device in real time; Step S22: The video data stream of the head-mounted high-resolution camera is obtained using FFmpeg technology; Step S23: The input video stream is opened through avformat_open_input, and the video stream parameters contained in the container format are parsed; Step S24: The stream metadata is obtained by calling avformat_find_stream_info, and the encoding context of each stream is determined; Step S25: The original data packet is extracted from the input video stream by calling av_read_frame in a loop, and is separated and stored in the corresponding queue according to the stream type; Step S26: The separated audio Packet sequence and video Packet sequence carry timestamp and stream index information, and create a codec context for each stream, and initialize the encoding parameters through avcodec_parameters_to_context; Step S27: The H.264 encoder is obtained by calling avcodec_find_decoder, and the decoder is opened through avcodec_open2; Step S28: Cycle Packet into decoder: input compressed data through avcodec_send_packet, call avcodec_receive_frame to output raw data frame; Step S29: Push the data frame decoded by FFmpeg to the wrist display terminal, and the internal operator can view the real-time video picture of the head-mounted high-resolution camera; Step S210: Push the data frame collected by the hole exploration device to the wrist display terminal, and the internal operator can view the real-time video picture inside the restricted space; Step S211: The measurement results of the step difference gap detection device are temporarily stored in the edge detection device memory in the form of a string.
[0009] Preferably, the hole exploration device is used to view the physically inaccessible area in the restricted space, which includes a flexible probe that can extend into the restricted space and an image acquisition device integrated at the front end of the probe; the outer diameter of the probe of the hole exploration device is not greater than 10 mm, and a flexible guide pipe and an illumination assembly are integrated inside; the guide pipe can realize positive and negative 90° bending turning to adapt to complex space structures; the illumination assembly uses high-brightness LED light source and configures a diffusion lens to realize uniform illumination; the image acquisition device integrated at the front end of the probe carries a miniature camera with a resolution of 1920x1080, supporting 30 frames per second dynamic video acquisition; the step difference gap detection device is a handheld high-precision non-contact optical measurement component, suitable for high-precision detection of step difference and gap of various size range gaps, with a measurement error less than or equal to ±25μm, and a measurement range of 0.1mm-25mm; the headset is used for voice interaction between the internal operator and the external data center, and the input and output of audio are realized through Windows Multimedia API.
[0010] Preferably, in step S3, the following steps are included: Step S31: The edge detection device and the external data processing center are both equipped with Ethernet gigabit network cards, supporting TCP / IP protocol stack, and the network communication function is realized through C++; Step S32: The external data processing center integrates WinSock network programming interface, including user interface module, data processing module and network communication module; Step S33: The edge detection device and the external data processing center interact data through gigabit transmission rate network cable, including data receiving and sending; Step S34: According to different types of data collected by internal data acquisition equipment, the type of data receiving and sending is determined, including transmission of headset video stream, transmission of hole exploration device video stream, transmission of step difference gap detection device string and transmission of headset audio; Step S35: The external data processing center creates a TCP socket based on WinSock: socket(), binds the specified IP address and port: bind(), and listens for connection requests: listen(). Step S36: The external data processing center receives the edge detection device connection through accept(), establishes a Socket handle mapping table, and supports multiple edge devices concurrent communication. Step S37: The external data processing center cyclically calls recv() to read the buffer data, parses the data packet according to the protocol, and extracts the encoded data after verification. Step S38: The external data processing center encapsulates the to-be-sent string according to the protocol, including adding the start symbol, length, check bit, and end symbol, and sends it to the target edge device through send(). Step S39: The edge detection device creates a TCP socket and connects the external data processing center IP address and port through connect(). Step S310: The edge detection device encapsulates the to-be-sent data according to the protocol and uploads it to the external data processing center through send(). When receiving instructions from the external data processing center, it parses the data packet and performs corresponding operations. Step S311: Set up a heartbeat packet mechanism to periodically send PING strings to detect network connection status and automatically reconnect in case of exception. Step S312: The internal and external network intercommunication enables the external data processing center to receive the head-mounted camera video stream, hole detection device video stream, step difference gap detection device string, and head-mounted earphone audio sent by the edge detection device in real time.
[0011] Preferably, in step S4, the following steps are included: Step S41: The external data processing center classifies the received head-mounted earphone video stream, hole detection device video stream, step difference gap detection device string, and head-mounted earphone audio and performs different operations. Step S42: For the head-mounted earphone video stream and hole detection device video stream, perform scratch detection and recognition based on the YOLOv8 network. Introduce the CBAM attention mechanism in the Neck part of YOLOv8, and use the SIoU loss function for parameter update. Step S43: Scale the image size to be detected to 640x640 and send it to the trained neural network model for detection. Step S44: The image first comes to the Backbone part in the model, passes through the P1 layer of Backbone, a standard 3x3 Conv module, and a convolution kernel size of 32, to obtain a feature map with a channel number of 32. Step S45: The feature map passes through the standard 3x3 Conv module of the P2 layer and the second layer C2f module of the network, and the number of feature map channels is 64. The C2f uses more jump connections and feature splicing to make the features more rich; Step S46: The feature map passes through the standard 3x3 Conv module of the P3 layer and the fourth layer C2f module of the network, and the number of channels is 128. The C2f is used for further fusion of the feature map; Step S47: The feature map passes through the standard 3x3 Conv module of the P4 layer and the sixth layer C2f module of the network, and the number of channels is 256. The C2f is used for further fusion of the feature map; Step S48: The feature map passes through the standard 3x3 Conv module of the P5 layer and the eighth layer C2f module of the network, and the number of channels is 512. The C2f is used for further fusion of the feature map; Step S49: The feature map passes through the SPPF module, and the number of channels is 512, which is used for spatial pyramid pooling to realize feature extraction of different scales and capture detailed information of the target at different scales; Step S410: The feature map comes to the Neck module in the model. First, the feature map is expanded by one time through upsampling operation and spliced with the feature map passing through the sixth layer. Then, the spliced feature passes through the twelfth layer C2f module of the network, and then passes through the CBAM attention mechanism module to improve the feature representation ability and strengthen the spatial information. The output dimension size of the feature map is 256; Step S411: The feature map of the previous layer is expanded by one time through upsampling and spliced with the feature map passing through the fourth layer. Then, the spliced feature passes through the sixteenth layer C2f module of the network, and then passes through the CBAM attention mechanism module. The output dimension size of the feature map is 128; Step S412: The feature map of the previous layer passes through the standard 3x3 Conv module and is spliced with the feature map passing through the thirteenth layer. Then, the spliced feature passes through the twentieth layer C2f module of the network, and then passes through the CBAM attention mechanism module. The output dimension size of the feature map is 256; Step S413: The feature map of the previous layer passes through the standard 3x3 Conv module and is spliced with the feature map passing through the ninth layer. Then, the spliced feature passes through the twenty-fourth layer C2f module of the network, and then passes through the CBAM attention mechanism module. The output dimension size of the feature map is 512; Step S414: The feature map passing through the seventeenth layer is input into the Head module of the network for identification and position positioning of the scratch target; Step S415: The feature map passing through the twenty-first layer is input into the Head module of the network for identification and position positioning of the scratch target; Step S416: input the feature map passing through the 25th layer to the Head module of the network to identify and locate the scratch target; Step S417: automatically record the scratch key frames detected by the head-mounted camera video stream and the hole exploration device video stream in the external data processing center database, and real-time view and trace historical data in the visual interface; Step S418: for the step difference gap detection device string data, it is segmented into detection categories and real values by the external data processing center, and the segmented data is automatically recorded in the external data processing center database according to the detection time; Step S419: the external data processing center is equipped with a professional database visual interactive interface, and the operation personnel can perform data adding, deleting, modifying and querying operations on the database based on the interface; Step S420: for the historical data in the database, representative data is automatically generated to generate a personalized report, and the report function module has a perfect comment embedding mechanism; Step S421: the external data processing center and the internal edge device are both equipped with a headset device, so as to build an internal and external real-time voice communication link to ensure that the internal and external can realize instant and efficient voice information interaction.
[0012] Preferably, in step S42, the calculation formula of the SIoU loss function is as follows: ; Wherein, represents the intersection over union between the predicted box and the real box, Δ represents the angle loss, and Ω represents the distance loss and the aspect ratio loss; The calculation formula of the angle loss Δ is as follows: ; Wherein, represents the Euclidean distance between the center points of the predicted box and the real box, represents the diagonal length of the minimum circumscribed rectangle of the predicted box and the real box; The calculation formula of the distance loss and the aspect ratio loss Ω is as follows: ; Wherein, and respectively represent the distance of the predicted box and the real box in the width and height directions, and respectively represent the width and height of the real box; In the training process, the parameters of the model are updated by minimizing the SIoU loss function.
[0013] Preferably, in step S42, the overall architecture of YOLOv8 includes three main parts: Backbone, Neck and Head; Backbone is responsible for extracting features from the input image, Neck fuses and enhances the features, and Head is used to predict the category and position of the target.
[0014] Preferably, the CBAM attention mechanism module includes a channel attention module and a spatial attention module.
[0015] Preferably, the channel attention module operates on the channel dimension of the feature map and assigns different weights to each channel to highlight important channel features. First, the input feature map Perform global average pooling and global maximum pooling respectively to obtain two global feature vectors and , which is calculated as follows: ; ; in, Representation feature map F The values of all channels in row i and column j; is the height of the feature map in the spatial dimension; is the width of the feature map in the spatial dimension; The two global feature vectors are input into a shared multi-layer perceptron respectively, and after the ReLU activation function and the Sigmoid activation function, the channel attention map M is obtained. c ∈R C : ; Among them, σ represents the Sigmoid activation function; v avg It is the global feature vector obtained by performing global average pooling on the input feature map F; v max It is the global feature vector obtained by performing global maximum pooling (GMP) on the input feature map F; Finally, the channel attention map M c With the input feature map F Perform channel-by-channel multiplication to obtain a feature map enhanced by channel attention : ; Among them, M c and F Element-wise multiplication.
[0016] Preferably, the spatial attention module then operates on the spatial dimension of the feature map, assigns different weights to each spatial position to highlight important spatial regions, and first performs channel attention enhancement on the feature map respectively, average pooling and maximum pooling are performed on the channel dimension to obtain two spatial feature maps and The calculation formula is as follows: ; ; The two spatial feature maps are spliced in the channel dimension to obtain a new feature map , and a 7x7 convolution layer and a Sigmoid activation function are used to obtain a spatial attention map : ; wherein, indicates the splicing operation in the channel dimension; Finally, the spatial attention map is multiplied element by element with the feature map enhanced by the channel attention to obtain the final feature map enhanced by the CBAM attention mechanism : ; wherein, and are multiplied element by element.
[0017] Compared with the prior art, the present application has the following advantages and beneficial effects: 1. The present application provides a limited space quality detection method based on multi-modal edge cloud transmission, which innovatively integrates a head-mounted camera, a flexible hole detection device and a step gap detection device to construct a millimeter-level three-dimensional space perception network, covers reachable and unreachable areas of personnel, and solves the problem of limited field of view in traditional manual detection. The edge detection device serves as a data transfer station, processes multi-modal data (video stream, measurement data, voice) in real time, and realizes high-resolution video local decoding and real-time display on a wrist terminal through FFmpeg technology.
[0018] Secondly, the application provides a limited space quality detection method based on multi-modal edge cloud transmission, which adopts a 10-gigabit Ethernet TCP / IP protocol stack, supports real-time transmission of high-resolution 30FPS video streams, automatically detects network abnormalities and reconnects through a heartbeat packet mechanism, solves the transmission delay problem caused by signal attenuation in a limited space, and guarantees that the video stream is not stuck. Differentiated transmission protocols are designed for video streams (head-mounted cameras / hole exploration devices), measurement strings (step difference gap data), and voice data to achieve efficient data packaging and analysis, and support concurrent communication of multiple edge detection devices (Socket handle mapping table).
[0019] Thirdly, the application provides a limited space quality detection method based on multi-modal edge cloud transmission, which introduces a CBAM attention mechanism in the Neck part of YOLOv8 to enhance scratch feature representation through channel and spatial attention weighting; a SIoU loss function is used instead of the traditional IoU to comprehensively consider the overlapping area, center point distance, and width-height ratio of the predicted frame and the real frame, significantly improving the micro-defect detection rate. Detection data (key frames, measurement values) are automatically entered into the database according to time and category, supporting visual traceability, multi-database switching, and personalized report generation (including comment embedding mechanism), realizing traceability and decision support of detection data. BRIEF DESCRIPTION OF DRAWINGS Figure 1 It is a flowchart of the application; Figure 2 It is a schematic diagram of the communication between the internal edge detection device and each data acquisition device in the application; Figure 3 It is a schematic diagram of high-resolution image encoding and decoding based on Ffmpeg in the application; Figure 4 It is a schematic diagram of TCP / IP protocol differentiated transmission of various data in the application; Figure 5 It is an improved YOLOv8 network model diagram in the application; Figure 6 It is a CBAM attention mechanism diagram in the application; Figure 7 It is a SIoU loss function diagram in the application; Figure 8 It is a database framework diagram in the application. DETAILED DESCRIPTION
[0020] The application will be further described in conjunction with the embodiments, but the implementation of the application is not limited thereto.
[0021] Example 1 This embodiment provides a limited space quality detection method based on multi-modal edge cloud transmission, which includes the following steps: Step S1: An internal operator carries an edge detection device into the workpiece to be measured and starts each data acquisition device in the confined space; Step S2: The wrist display terminal worn by the internal operator displays the video images of the head-mounted camera and the borescope equipment in real time; Step S3: The video data, step gap detection data, and real-time voice interaction data from the head-mounted camera and borescope equipment inside the confined space are transmitted to an external data processing center via TCP / IP; Step S4: the external data processing center performs scratch detection on the received video data, and selectively enters the scratch detection key frame data, video stream data, and step gap detection data into a database.
[0022] Example 2 This embodiment provides a confined space quality detection method based on multimodal edge cloud transmission, including the following steps: Step S1: An internal operator carries an edge detection device into the workpiece to be measured and starts each data acquisition device in the confined space; Step S2: The wrist display terminal worn by the internal operator displays the video images of the head-mounted camera and the borescope equipment in real time; Step S3: The video data, step gap detection data, and real-time voice interaction data from the head-mounted camera and borescope equipment inside the confined space are transmitted to an external data processing center via TCP / IP; Step S4: the external data processing center performs scratch detection on the received video data, and selectively enters the scratch detection key frame data, video stream data, and step gap detection data into a database.
[0023] Among them, in step S1, the workpiece to be inspected is determined, and the internal operator carries various data acquisition devices and edge detection devices into the narrow space of the workpiece to be inspected; each data acquisition device includes a head-mounted camera, a borescope device, a step gap detection device and a headset; the edge detection device is connected to each data acquisition device as a data transfer station, and the data acquisition device is used to scan the inside of the workpiece and transmit the data to the edge detection device; each set of portable equipment is equipped with a lighting device, and the operator turns on or off the lighting device as needed to ensure the availability of the collected data.
[0024] Wherein, the step S2 includes the following steps: Step S21: The edge detection device is equipped with a wrist-mounted small LCD screen, and internal staff can view the video images of the head-mounted camera and the borescope device in real time; Step S22: using FFmpeg technology to obtain the video data stream of the head-mounted high-resolution camera; Step S23: open the input video stream by avformat_open_input, and parse the video stream parameters contained in the container format; Step S24: call avformat_find_stream_info to obtain stream metadata, and determine the encoding context of each stream; Step S25: call av_read_frame in a loop to extract raw data packets from the input video stream, separate them by stream type, and store them in the corresponding queue; Step S26: separate the audio Packet sequence and the video Packet sequence, carry the timestamp and stream index information, create a codec context for each stream, and initialize the encoding parameters by avcodec_parameters_to_context; Step S27: call avcodec_find_decoder to obtain the H.264 encoder, and open the decoder by avcodec_open2; Step S28: loop to send Packets to the decoder: input compressed data by avcodec_send_packet, and output raw data frames by calling avcodec_receive_frame; Step S29: push the data frames decoded by FFmpeg to the wrist display terminal, and the internal operator can view the real-time video picture of the head-mounted high-resolution camera; Step S210: push the data frames collected by the hole exploration device to the wrist display terminal, and the internal operator can view the real-time video picture inside the restricted space; Step S211: the measurement results of the step difference gap detection device are temporarily stored in the edge detection device memory in the form of a string.
[0025] The hole exploration device is used to view physically inaccessible areas in the restricted space, and includes a flexible probe that can extend into the restricted space and an image acquisition device integrated at the front end of the probe. The outer diameter of the probe of the hole exploration device is not greater than 10 mm, and a flexible guide pipe and a lighting assembly are integrated inside. The guide pipe can realize positive and negative 90° bending turning, and is suitable for complex space structures. The lighting assembly uses a high-brightness LED light source, and is configured with a diffusion lens to realize uniform illumination. The image acquisition device integrated at the front end of the probe is equipped with a miniature camera with a resolution of 1920x1080, and supports 30 frames per second dynamic video acquisition. The step difference gap detection device is a handheld high-precision non-contact optical measurement component, which is suitable for high-precision detection of step differences and gaps of various size range gaps, with a measurement error less than or equal to ±25μm, and a measurement range of 0.1mm-25mm. The headset is used for voice interaction between the internal operator and the external data center, and realizes audio input and output through Windows Multimedia API.
[0026] Wherein, the step S3 includes the following steps: Step S31: The edge detection device and the external data processing center are both equipped with Ethernet 10G network cards, support TCP / IP protocol stack, and implement network communication functions through C++; Step S32: The external data processing center integrates the WinSock network programming interface, including a user interface module, a data processing module, and a network communication module; Step S33: The edge detection device and the external data processing center perform data exchange via a 10G transmission rate network cable, including data reception and transmission; Step S34: The types of data received and sent are determined based on the different types of data collected by the internal data acquisition devices, including the transmission of the headphone video stream, the transmission of the borescope device video stream, the transmission of the step gap detection device character string, and the headphone audio transmission; Step S35: the external data processing center creates a TCP socket based on WinSock: socket(), binds the specified IP address and port: bind(), and listens for connection requests: listen(); Step S36: The external data processing center receives the edge detection device connection through accept(), establishes a socket handle mapping table, and supports concurrent communication of multiple edge devices; Step S37: The external data processing center cyclically calls recv() to read the buffer data, parses the data packet according to the protocol, and extracts the encoded data after verification; Step S38: The external data processing center encapsulates the string to be sent according to the protocol, including adding the start character, length, check bit, and end character, and sends it to the target edge device through send(); Step S39: The edge detection device creates a TCP socket and connects to the external data processing center IP address and port through connect(); Step S310: The edge detection device encapsulates the data to be sent according to the protocol and uploads it to the external data processing center through send(); when receiving instructions from the external data processing center, it parses the data packet and performs corresponding operations; Step S311: Set up a heartbeat packet mechanism to periodically send a PING string to detect the network connection status and automatically reconnect when an abnormality occurs; Step S312: The internal and external networks are interconnected so that the external data processing center receives in real time the head-mounted camera video stream, borescope device video stream, step gap detection device character string and headphone audio sent by the edge detection device.
[0027] Wherein, the step S4 includes the following steps: Step S41: The external data processing center makes different operations on the received headset video stream, hole probe device video stream, step gap detection device string and headset audio classification; Step S42: For the headset video stream and the hole probe device video stream, scratch detection and recognition are performed based on the YOLOv8 network, the CBAM attention mechanism is introduced in the Neck part of YOLOv8, and the SIoU loss function is used for parameter updating; Step S43: The image to be detected is scaled to 640x640 and sent to the trained neural network model for detection; Step S44: The image first comes to the Backbone part in the model, passes through the P1 layer of Backbone, a standard 3x3 Conv module, and the convolution kernel size is 32, so as to obtain a feature map with a channel number of 32; Step S45: The feature map passes through the standard 3x3 Conv module of P2 layer and the 2nd layer C2f module of the network, the channel number of the feature map is 64, and C2f adopts more jump connections and feature splicing to make the features more rich; Step S46: The feature map passes through the standard 3x3 Conv module of P3 layer and the 4th layer C2f module of the network, the channel number is 128, and C2f is used for further fusion of the feature map; Step S47: The feature map passes through the standard 3x3 Conv module of P4 layer and the 6th layer C2f module of the network, the channel number is 256, and C2f is used for further fusion of the feature map; Step S48: The feature map passes through the standard 3x3 Conv module of P5 layer and the 8th layer C2f module of the network, the channel number is 512, and C2f is used for further fusion of the feature map; Step S49: The feature map passes through the SPPF module, the channel number is 512, which is used for spatial pyramid pooling to realize feature extraction of different scales and capture detailed information of the target at different scales; Step S410: The feature map comes to the Neck module in the model, first expands the feature map by one time through upsampling operation, splices with the feature map passing through the 6th layer, then passes the spliced feature through the 12th layer C2f module of the network, and then passes through the CBAM attention mechanism module to improve the feature representation ability and strengthen the spatial information, and outputs a feature map with a dimension size of 256; Step S411: The feature map of the last layer is expanded by one time through upsampling, spliced with the feature map passing through the 4th layer, then the spliced feature passes through the 16th layer C2f module of the network, and then passes through the CBAM attention mechanism module, and outputs a feature map with a dimension size of 128; Step S412: The feature map of the previous layer is passed through a standard 3x3 Conv module, spliced with the feature map passing through the 13th layer, and then the spliced features are passed through the 20th layer C2f module of the network, followed by a CBAM attention mechanism module, outputting a feature map with a dimension size of 256; Step S413: The feature map of the previous layer is passed through a standard 3x3 Conv module, spliced with the feature map passing through the 9th layer, and then the spliced features are passed through the 24th layer C2f module of the network, followed by a CBAM attention mechanism module, outputting a feature map with a dimension size of 512; Step S414: The feature map passing through the 17th layer is input into the Head module of the network for scratch target recognition and position positioning; Step S415: The feature map passing through the 21st layer is input into the Head module of the network for scratch target recognition and position positioning; Step S416: The feature map passing through the 25th layer is input into the Head module of the network for scratch target recognition and position positioning; Step S417: The scratch key frame detected by the head-mounted camera video stream and the hole exploration device video stream is automatically entered into the external data processing center database, and historical data can be viewed and traced in real time in the visualization interface; Step S418: For the step difference gap detection device string data, it is segmented into detection categories and real values by the external data processing center, and the segmented data is automatically entered into the external data processing center database according to the detection time; Step S419: The external data processing center is equipped with a professional database visualization interface, and the operation personnel can perform data adding, deleting, modifying and querying operations on the database based on the interface; Step S420: For historical data in the database, representative data is automatically generated into a personalized report, and the report function module has a perfect comment embedding mechanism; Step S421: The external data processing center and the internal edge device are both equipped with a headset device, which builds a real-time voice communication link between the inside and the outside, ensuring that the inside and the outside can realize instant and efficient voice information interaction.
[0028] In step S42, the calculation formula of the SIoU loss function is as follows: ; Wherein, represents the intersection over union between the predicted box and the real box, Δ represents the angle loss, and Ω represents the distance loss and the aspect ratio loss; The calculation formula of the angle loss Δ is as follows: ; wherein, denotes the Euclidean distance between the center points of the predicted box and the real box, denotes the length of the diagonal of the minimum circumscribed rectangle of the predicted box and the real box; The calculation formula of the distance loss and the aspect ratio loss Ω is as follows: ; wherein, and denote the distance of the predicted box and the real box in the width and height directions, and denote the width and height of the real box, respectively; In the training process, the parameters of the model are updated by minimizing the SIoU loss function.
[0029] wherein, in the step S42, the overall architecture of YOLOv8 includes three main parts of Backbone, Neck and Head; Backbone is responsible for extracting features from input images, Neck fuses and enhances the features, and Head is used to predict the category and position of the target.
[0030] wherein, the CBAM attention mechanism module includes a channel attention module and a spatial attention module.
[0031] wherein, the channel attention module assigns different weights to each channel by operating on the channel dimension of the feature map, in order to highlight important channel features; first, the input feature map is respectively subjected to global average pooling and global maximum pooling to obtain two global feature vectors and , whose calculation formula is as follows: ; ; wherein, denotes the value of all channels of the feature map F in the i-th row and j-th column; is the height of the feature map in the spatial dimension; is the width of the feature map in the spatial dimension; used to describe the spatial size of the input feature map , wherein C is the number of channels.
[0032] The two global feature vectors are respectively input into a shared multi-layer perceptron, and after ReLU activation function and Sigmoid activation function, the channel attention map M c ∈R C : ; Among them, σ represents the Sigmoid activation function; v avg It is the global feature vector obtained by performing global average pooling on the input feature map F; v max It is the global feature vector obtained by performing global maximum pooling (GMP) on the input feature map F; Finally, the channel attention map M c With the input feature map F Perform channel-by-channel multiplication to obtain a feature map enhanced by channel attention : ; Among them, M c and F Element-wise multiplication.
[0033] The spatial attention module operates on the spatial dimension of the feature map and assigns different weights to each spatial position to highlight the important spatial area. First, the feature map enhanced by channel attention is Perform average pooling and maximum pooling on the channel dimension respectively to obtain two spatial feature maps and , which is calculated as follows: ; ; The two spatial feature maps are spliced in the channel dimension to obtain a new feature map , and obtain the spatial attention map through a 7×7 convolution layer and Sigmoid activation function : ; in, Represents the concatenation operation in the channel dimension; Finally, the spatial attention map Feature map enhanced by channel attention Perform element-by-element multiplication to obtain the final feature map enhanced by the CBAM attention mechanism : ; in, and Element-wise multiplication.
[0034] Example 3 like Figure 1 As shown, the embodiment of the present application provides a confined space quality detection method based on multimodal edge cloud transmission, including: S1: The internal operator carries the edge detection device into the workpiece to be tested, and starts the data acquisition device in the restricted space.
[0035] S1 includes the following steps: S11: Determine the workpiece to be tested, and carry the data acquisition device and the edge device into the narrow space of the workpiece to be tested by the internal operator.
[0036] S12: The edge device is connected to each data acquisition device as a data transfer station, and uses the data acquisition device to scan the inside of the workpiece and transmit it to the edge device. (As shown in the internal edge device and each data acquisition device communication diagram) Figure 2 S13: The internal data acquisition device includes a head-mounted camera, a hole probe device, a step gap detection device, and a head-mounted earphone.
[0037] S14: In order to avoid the problem of difficulty in work due to insufficient light inside the workpiece, each set of portable device is equipped with a lighting device, and the operator can turn on or off the lighting device according to the need to ensure the availability of the collected data.
[0038] S2: The wrist display terminal worn by the internal operator displays the video pictures of the head-mounted camera and the hole probe device in real time.
[0039] S2 includes the following steps: S21: The edge device is equipped with a wrist small liquid crystal display screen, and the internal operator can view the video pictures of the head-mounted camera and the hole probe device in real time.
[0040] S22: Use FFmpeg technology to obtain the video data stream of the head-mounted high-resolution camera. (As shown in the Ffmpeg high-resolution image encoding and decoding principle diagram) Figure 3 S23: Open the input video stream by avformat_open_input, and parse the video stream parameters contained in the container format.
[0041] S24: Call avformat_find_stream_info to obtain stream metadata and determine the encoding context of each stream.
[0042] S25: Recursively call av_read_frame to extract raw data packets from the input video stream, separate them by stream type, and store them in the corresponding queue.
[0043] S26: Separate the audio Packet sequence and the video Packet sequence, carry the timestamp and stream index information, create a codec context for each stream, and initialize the encoding parameters by avcodec_parameters_to_context.
[0044] S27: Call avcodec_find_decoder to get the H.264 encoder, and open the decoder through avcodec_open2.
[0045] S28: Loop to send Packet into the decoder: input compressed data through avcodec_send_packet, and call avcodec_receive_frame to output raw data frame.
[0046] S29: Push the data frame decoded by FFmpeg to the wrist display terminal, and internal workers can view the real-time video picture of the head-mounted high-resolution camera.
[0047] S210: The hole exploration device is used to view the physically inaccessible area of the personnel in the restricted space, which includes a flexible probe that can extend into the restricted space and an image acquisition device integrated at the front end of the probe.
[0048] S211: The outer diameter of the probe is not greater than 10mm, and a flexible catheter and an illumination assembly are integrated inside, the catheter can realize positive and negative 90° bending turning to adapt to complex space structure; the illumination assembly uses high-brightness LED light source, and is configured with a diffusion lens to realize uniform illumination and avoid strong light reflection.
[0049] S212: The image acquisition device integrated at the front end of the probe carries a miniature camera with a resolution of 1920x1080, supporting 30 frames per second dynamic video acquisition.
[0050] S213: Push the data frame collected by the hole exploration device to the wrist display terminal, and internal workers can view the real-time video picture inside the restricted space.
[0051] S214: The step difference gap detection device is a handheld high-precision non-contact optical measurement component, suitable for high-precision detection of step difference and gap of various size range gaps, with a measurement error less than or equal to ±25μm, and a measurement range of 0.1mm-25mm.
[0052] S215: The measurement results of the step difference gap detection device are temporarily stored in the edge device memory in the form of a string.
[0053] S216: The headset is used for voice interaction between internal workers and external data center, and the input and output of audio are realized through Windows Multimedia API.
[0054] Step S3: The video data of the head-mounted camera and the hole exploration device, the step difference gap detection data, and the real-time voice interaction data in the restricted space are transmitted to the external data processing center through TCP / IP. Wherein, step S3 comprises: S31: The internal edge device and the external data processing center are both equipped with Ethernet gigabit network cards, support TCP / IP protocol stack, and realize network communication function through C++.
[0055] S32: The external data processing center integrates WinSock network programming interface, contains user interface module, data processing module and network communication module.
[0056] S33: The internal edge device and the external data processing center interact data through gigabit transmission rate network cable, including data receiving and sending.
[0057] S34: According to different types of data collected by internal data collection device, the type of data receiving and sending is determined, including transmission of headset video stream, transmission of hole detection device video stream, transmission of step difference gap detection device string and transmission of headset audio. (As shown in the TCP / IP protocol differentiated transmission of various data schematic diagram) Figure 4 S35: The external data processing center creates TCP socket based on WinSock: socket(), binds specified IP address and port: bind(), and listens to connection request: listen().
[0058] S36: The external data processing center receives internal edge device connection through accept(), establishes socket handle mapping table, and supports multi-edge device concurrent communication.
[0059] S37: The external data processing center cyclically calls recv() to read buffer data, parses data packet according to protocol, extracts encoded data after verification.
[0060] S38: The external data processing center encapsulates (adds start symbol, length, check bit and end symbol) the to-be-sent string according to protocol, and sends it to the target edge device through send().
[0061] S39: The internal edge device creates TCP socket, and connects external data processing center IP address and port through connect().
[0062] S310: The internal edge device encapsulates to-be-sent data according to protocol, and uploads it to the external data processing center through send(). When receiving external data processing center instruction, the data packet is parsed and corresponding operation is executed.
[0063] S311: Heartbeat packet mechanism (periodically send PING string) is set to detect network connection state, and automatically reconnect when abnormal.
[0064] S312: The internal and external network intercommunication enables the external data processing center to receive the head-mounted camera video stream, the hole probe device video stream, the step difference gap detection device string and the headset audio sent by the internal edge device in real time.
[0065] Step S4: The external data processing center performs scratch detection on the received video data, and selectively enters the scratch detection key frame data, the video stream data and the step difference gap detection data into the database.
[0066] Among them, step S4 includes: S41: The external data processing center classifies the received headset video stream, hole probe device video stream, step difference gap detection device string and headset audio to make different operations; S42: For the headset video stream and the hole probe device video stream, YOLOv8 network is used for scratch detection and identification, CBAM (Convolutional Block Attention Module) attention mechanism is introduced in the Neck part of YOLOv8, and SIoU (Scale-Invariant Intersection over Union) loss function is used for parameter update, which significantly improves the performance of scratch detection. (As shown in the improved YOLOv8 network model diagram Figure 5 S43: The overall architecture of YOLOv8 includes three main parts: Backbone, Neck and Head. Backbone is responsible for extracting features from input images, Neck fuses and enhances features, and Head is used to predict the category and position of the target.
[0067] S44: The image size to be detected is scaled to 640x640 and sent to the trained neural network model for detection; S45: The image first comes to the Backbone part in the model, and after the P1 layer of Backbone, a standard 3x3 Conv module with a convolution kernel size of 32, a feature map with a channel number of 32 is obtained; S46: The feature map passes through the standard 3x3 Conv module of P2 layer and the 2nd layer C2f module of the network, and the channel number of the feature map is 64. C2f uses more jump connections and feature splicing to make the features more rich; S47: The feature map passes through the standard 3x3 Conv module of P3 layer and the 4th layer C2f module of the network, and the channel number is 128. C2f is used to further fuse the feature map; S48: The feature map passes through the standard 3*3 Conv module of the P4 layer and the C2f module of the 6th layer of the network, the channel number is 256, and the C2f is used for further fusion of the feature map; S49: The feature map passes through the standard 3*3 Conv module of the P5 layer and the C2f module of the 8th layer of the network, the channel number is 512, and the C2f is used for further fusion of the feature map; S410: The feature map passes through the SPPF module, the channel number is 512, which is used for spatial pyramid pooling, realizing feature extraction of different scales and capturing detailed information of the target at different scales; S411: The feature map comes to the Neck module in the model, first passes through the upsampling operation to expand the feature map by one time, is spliced with the feature map passing through the 6th layer, and then the spliced feature passes through the C2f module of the 12th layer of the network, and then passes through the CBAM attention mechanism module, which is used for improving the feature representation ability and strengthening the spatial information, and outputs a feature map with a dimension of 256; S412: The feature map of the previous layer is expanded by one time through upsampling, and is spliced with the feature map passing through the 4th layer, and then the spliced feature passes through the C2f module of the 16th layer of the network, and then passes through the CBAM attention mechanism module, and outputs a feature map with a dimension of 128; (as Figure 6 The CBAM attention mechanism principle diagram is shown) S413: In order to enhance the attention of YOLOv8 to the scratch feature, the CBAM attention mechanism is introduced in the Neck part of the present application. CBAM is composed of a channel attention module (Channel Attention Module) and a spatial attention module (Spatial Attention Module).
[0068] S414: The channel attention module operates on the channel dimension of the feature map, and assigns different weights to each channel to highlight important channel features. Specifically, the channel attention module first performs global average pooling (Global Average Pooling, GAP) and global maximum pooling (Global Max Pooling, GMP) on the input feature map respectively, to obtain two global feature vectors and The calculation formula is as follows: ; ; Among them, represents the value of all channels of the feature map F in the i-th row and j-th column; is the height of the feature map in the spatial dimension; is the width of the feature map in spatial dimension; used to describe the spatial size of the input feature map , where C is the number of channels.
[0069] The two global feature vectors are input into a shared Multilayer Perceptron (MLP) respectively, and after ReLU activation function and Sigmoid activation function, a channel attention map M is obtained: ; wherein, σ represents the Sigmoid activation function; v avg is a global feature vector obtained by global average pooling (GAP) on the input feature map F; v max is a global feature vector obtained by global maximum pooling (GMP) on the input feature map F.
[0070] Finally, the channel attention map M c is multiplied with the input feature map F F channel by channel, to obtain a feature map enhanced by channel attention M : ; wherein, M c is multiplied with F element by element.
[0071] S415: The spatial attention module then operates on the spatial dimension of the feature map, and assigns different weights to each spatial position to highlight important spatial regions. The spatial attention module first performs average pooling and maximum pooling on the channel dimension of the feature map enhanced by channel attention M , respectively, to obtain two spatial feature maps and , whose calculation formulas are as follows: ; ; The two spatial feature maps are spliced in the channel dimension to obtain a new feature map , and through a 7x7 convolution layer and a Sigmoid activation function, a spatial attention map M is obtained: ; wherein, represents the splicing operation in the channel dimension; Finally, the spatial attention map M is multiplied with the feature map enhanced by channel attention M Element-wise multiplication is performed to obtain the final feature map enhanced by the CBAM attention mechanism : .
[0072] wherein, and element-wise multiplication.
[0073] S416: The feature map of the previous layer is passed through a standard 3*3 Conv module, spliced with the feature map passed through the 13th layer, and then the spliced feature is passed through the 20th layer C2f module of the network, followed by the CBAM attention mechanism module, and a feature map with a dimension size of 256 is output; S417: The feature map of the previous layer is passed through a standard 3*3 Conv module, spliced with the feature map passed through the 9th layer, and then the spliced feature is passed through the 24th layer C2f module of the network, followed by the CBAM attention mechanism module, and a feature map with a dimension size of 512 is output; S418: The feature map passed through the 17th layer is input into the Head module of the network for scratch target recognition and position positioning; S419: The feature map passed through the 21st layer is input into the Head module of the network for scratch target recognition and position positioning; S420: The feature map passed through the 25th layer is input into the Head module of the network for scratch target recognition and position positioning; S421: In order to more accurately update the model parameters, the present application adopts SIoU as the loss function. The SIoU loss function comprehensively considers the overlapping area, center point distance and width-height ratio between the predicted frame and the real frame, and can more effectively optimize the target detection model. (As shown in the SIoU loss function principle diagram Figure 7 , wherein B is the predicted frame, B GT is the real frame, σ represents the distance between the center of the predicted frame and the center of the real frame, C h is the vertical distance between the center of the predicted frame and the center of the real frame, C w is the horizontal distance between the center of the predicted frame and the center of the real frame, and α and β represent arcsin(C h / σ) and arcsin(C w / σ), respectively.) The calculation formula of the SIoU loss function is as follows: ; wherein, represents the intersection over union (Intersection over Union) between the predicted frame and the real frame, Δ represents the angle loss, and Ω represents the distance loss and width-height ratio loss.
[0074] The calculation formula of the angle loss Δ is as follows: ; Wherein, represents the Euclidean distance between the center points of the prediction box and the real box, represents the diagonal length of the minimum circumscribed rectangle of the prediction box and the real box.
[0075] The calculation formula of the distance loss and the aspect ratio loss Ω is as follows: ; Wherein, and respectively represent the distance of the prediction box and the real box in the width and height directions, and respectively represent the width and height of the real box.
[0076] In the training process, the parameters of the model are updated by minimizing the SIoU loss function, so that the model can more accurately detect scratches.
[0077] S422: The scratch key frame detected by the head-mounted camera video stream and the hole detection equipment video stream is automatically entered into the external data processing center database, and the historical data can be viewed and traced in real time in the visualization interface. (As shown in the database framework diagram of Figure 8 ) S423: For the step difference gap detection equipment string data, it is segmented into detection categories (such as step difference, gap, chamfer, round corner, etc.) and real values by the external data processing center, and the segmented data is automatically entered into the external data processing center database according to the detection time.
[0078] S424: The external data processing center is equipped with a professional database visualization interface, and the operation personnel can perform data adding, deleting, modifying and querying operations on the database based on this interface. For different detection scenarios, independent databases can be constructed, and seamless switching between databases is supported, thereby realizing the full-process traceability and precise controllability of detection data.
[0079] S425: For the historical data in the database, representative data can be selected to automatically generate personalized reports. This report function module has a perfect comment embedding mechanism, allowing the operator to enter professional evaluations for the results of a certain detection task, and seamlessly integrating the comments into the corresponding report system.
[0080] S426: The external data processing center and the internal edge equipment are both equipped with a headset device, which builds a real-time voice communication link between the inside and the outside, ensuring that the inside and the outside can realize instant and efficient voice information interaction.
[0081] The above merely illustrates the preferred embodiments of the present application, but is not intended to limit the present application in any form. Any simple modification or equivalent change of the above embodiments according to the technical essence of the present application shall fall within the protection scope of the present application.
Claims
1. A method for limited space quality detection based on multi-modal edge cloud transmission, characterized in that, The method comprises the following steps: Step S1: An internal operator carries an edge detection device into a workpiece to be detected, and starts each data acquisition device in a restricted space; Step S2: A wrist-mounted display terminal worn by the internal operator displays video pictures of a head-mounted camera and a hole probe device in real time; Step S3: Video data of the head-mounted camera and the hole probe device, step difference gap detection data, and real-time voice interaction data in the restricted space are transmitted to an external data processing center through TCP / IP; Step S4: The external data processing center performs scratch detection on the received video data, and selectively enters scratch detection key frame data, video stream data, and step difference gap detection data into a database.
2. The method of claim 1, wherein the method is based on a multi-modal edge cloud transmission. In the step S1, the workpiece to be detected is determined, and each data acquisition device and the edge detection device are carried into a narrow space of the workpiece to be detected by the internal operator; the data acquisition device includes a head-mounted camera, a hole probe device, a step difference gap detection device, and a head-mounted earphone; the edge detection device is connected with each data acquisition device as a data transfer station, the workpiece is scanned by using the data acquisition device and is transmitted to the edge detection device; each set of portable device is provided with an illumination device, and the operator turns on or off the illumination device according to needs to ensure the usability of the collected data. 3.The method of claim 2, wherein: In the step S2, the following steps are included: Step S21: The edge detection device is provided with a wrist-mounted small liquid crystal display screen, and the internal operator views video pictures of the head-mounted camera and the hole probe device in real time; Step S22: Video data stream of the head-mounted high-resolution camera is obtained by using FFmpeg technology; Step S23: An input video stream is opened by using avformat_open_input, and video stream parameters contained in a container format are parsed; Step S24: Stream metadata is obtained by using avformat_find_stream_info, and coding contexts of each stream are determined; Step S25: Raw data packets are extracted from the input video stream by using av_read_frame, are separated according to stream types, and are stored in corresponding queues; Step S26: Separated audio Packet sequences and video Packet sequences carry timestamp and stream index information, coding and decoding contexts are created for each stream, and coding parameters are initialized by using avcodec_parameters_to_context; Step S27: An H.264 encoder is obtained by using avcodec_find_decoder, and the decoder is opened by using avcodec_open2; Step S28: Packets are sent into the decoder in a loop: compressed data is input by using avcodec_send_packet, and raw data frames are output by using avcodec_receive_frame; Step S29: Data frames decoded by using FFmpeg are pushed to the wrist-mounted display terminal, and the internal operator views real-time video pictures of the head-mounted high-resolution camera; Step S210: Data frames collected by the hole probe device are pushed to the wrist-mounted display terminal, and the internal operator can view real-time video pictures in the restricted space; Step S211: The measurement result of the step difference gap detection device is temporarily stored in the edge detection device memory in the form of a string.
4. The method of claim 3, wherein the method is based on a multi-modal edge cloud transmission. The hole exploration device is used for viewing physically inaccessible areas of personnel in a restricted space, and includes a flexible probe that can extend into the restricted space and an image acquisition device integrated at the front end of the probe; the outer diameter of the probe of the hole exploration device is not greater than 10 mm, and a flexible guide pipe and an illumination assembly are integrated inside; the guide pipe can realize positive and negative 90° bending turning to adapt to complex space structures; the illumination assembly adopts a high-brightness LED light source and is configured with a diffusion lens to realize uniform illumination; the image acquisition device integrated at the front end of the probe is equipped with a miniature camera with a resolution of 1920*1080 and supports 30 frames per second dynamic video acquisition; the step difference gap detection device is a handheld high-precision non-contact optical measurement component, which is suitable for high-precision detection of step differences and gaps of various size range gaps, and the measurement error is less than or equal to ±25μm, and the measurement range is 0.1mm-25mm; the headset is used for voice interaction between internal workers and external data center, and the input and output of audio are realized through Windows Multimedia API.
5. The method of claim 4, wherein: The step S3 includes the following steps: Step S31: The edge detection device and the external data processing center are both equipped with Ethernet gigabit network cards, support TCP / IP protocol stack, and realize network communication function through C++; Step S32: The external data processing center integrates WinSock network programming interface, including user interface module, data processing module and network communication module; Step S33: The edge detection device and the external data processing center interact data through gigabit transmission rate network cable, including data receiving and sending; Step S34: According to different types of data collected by the internal data acquisition device, the type of data receiving and sending is determined, including transmission of headset video stream, transmission of hole exploration device video stream, transmission of step difference gap detection device string and transmission of headset audio; Step S35: The external data processing center creates a TCP socket based on WinSock: socket(), binds a specified IP address and port: bind(), and listens to connection requests: listen(); Step S36: The external data processing center receives the edge detection device connection through accept(), establishes a socket handle mapping table, and supports multiple edge devices concurrent communication; Step S37: The external data processing center cyclically calls recv() to read buffer data, parses data packets according to the protocol, extracts encoded data after verification; Step S38: The external data processing center encapsulates the to-be-sent string according to the protocol, including adding start symbol, length, check bit and end symbol, and sends it to the target edge device through send(); Step S39: The edge detection device creates a TCP socket and connects the IP address and port of the external data processing center through connect(). Step S310: The edge detection device encapsulates the data to be sent according to the protocol, uploads to the external data processing center through send(), and parses the data packet and executes the corresponding operation when receiving the instruction of the external data processing center; Step S311: A heartbeat packet mechanism is set, a PING string is sent periodically to detect the network connection state, and the network is automatically reconnected in an abnormal state; Step S312: The internal and external network intercommunication enables the external data processing center to receive the head-mounted camera video stream, hole detection device video stream, step difference gap detection device string and head-mounted earphone audio sent by the edge detection device in real time.
6. The method of claim 5, wherein: The step S4 comprises the following steps: Step S41: The external data processing center classifies the received head-mounted earphone video stream, hole detection device video stream, step difference gap detection device string and head-mounted earphone audio to make different operations; Step S42: The head-mounted earphone video stream and hole detection device video stream are subjected to scratch detection and recognition based on the YOLOv8 network, the CBAM attention mechanism is introduced in the Neck part of YOLOv8, and the SIoU loss function is used for parameter updating; Step S43: The image to be detected is scaled to 640x640 and sent to the trained neural network model for detection; Step S44: The image first comes to the Backbone part in the model, passes through the P1 layer of Backbone, a standard 3x3 Conv module, and a convolution kernel size of 32, so as to obtain a feature map with a channel number of 32; Step S45: The feature map passes through the standard 3x3 Conv module of P2 layer and the 2nd layer C2f module of the network, the channel number of the feature map is 64, and C2f adopts more jump connection and feature splicing to make the feature more rich; Step S46: The feature map passes through the standard 3x3 Conv module of P3 layer and the 4th layer C2f module of the network, the channel number is 128, and C2f is used for further fusion of the feature map; Step S47: The feature map passes through the standard 3x3 Conv module of P4 layer and the 6th layer C2f module of the network, the channel number is 256, and C2f is used for further fusion of the feature map; Step S48: The feature map passes through the standard 3x3 Conv module of P5 layer and the 8th layer C2f module of the network, the channel number is 512, and C2f is used for further fusion of the feature map; Step S49: The feature map passes through the SPPF module, the channel number is 512, and is used for spatial pyramid pooling to realize feature extraction of different scales and capture detailed information of the target at different scales; Step S410: The feature map comes to the Neck module in the model, first expands the feature map by one time through upsampling operation, splices with the feature map of the 6th layer, then passes the spliced feature through the 12th layer C2f module of the network, and then passes through the CBAM attention mechanism module to improve the feature representation ability and strengthen the spatial information, and outputs a feature map with a dimension size of 256; Step S411: The feature map of the previous layer is expanded by one time through upsampling, spliced with the feature map passing through the 4th layer, and then the spliced feature passes through the 16th layer C2f module of the network, followed by the CBAM attention mechanism module, outputting a feature map with a dimension size of 128; Step S412: The feature map of the previous layer passes through a standard 3x3 Conv module, is spliced with the feature map passing through the 13th layer, and then the spliced feature passes through the 20th layer C2f module of the network, followed by the CBAM attention mechanism module, outputting a feature map with a dimension size of 256; Step S413: The feature map of the previous layer passes through a standard 3x3 Conv module, is spliced with the feature map passing through the 9th layer, and then the spliced feature passes through the 24th layer C2f module of the network, followed by the CBAM attention mechanism module, outputting a feature map with a dimension size of 512; Step S414: The feature map passing through the 17th layer is input into the Head module of the network for scratch target recognition and position positioning; Step S415: The feature map passing through the 21st layer is input into the Head module of the network for scratch target recognition and position positioning; Step S416: The feature map passing through the 25th layer is input into the Head module of the network for scratch target recognition and position positioning; Step S417: The scratch key frames detected by the head-mounted camera video stream and the hole probe device video stream are automatically entered into the external data processing center database, and the real-time viewing and historical data tracing can be realized in the visual interface; Step S418: For the step difference gap detection device string data, it is segmented into detection categories and real values by the external data processing center, and the segmented data is automatically entered into the external data processing center database according to the detection time; Step S419: The external data processing center is equipped with a professional database visual interactive interface, and the operation personnel can perform data adding, deleting, modifying and querying operations on the database based on the interface; Step S420: For the historical data in the database, representative data is automatically generated to generate a personalized report, and the report function module has a perfect comment embedding mechanism; Step S421: The external data processing center and the internal edge device are both equipped with a headset device, which builds a real-time voice communication link between the inside and the outside, ensuring that the inside and the outside can realize instant and efficient voice information interaction.
7. The method of claim 6, wherein: In step S42, the calculation formula of the SIoU loss function is as follows: ; wherein, denotes the intersection over union between the predicted and the true box, denotes the angle loss, and denotes the distance and aspect ratio loss. The calculation formula of the angle loss Δ is as follows: ; wherein, represents the Euclidean distance between the center points of the predicted and true bounding boxes, represents the length of the diagonal of the minimum circumscribed rectangle of the predicted and true bounding boxes; The calculation formula of the distance loss and the aspect ratio loss Ω is as follows: ; wherein, and respectively represent the distance of the predicted box and the real box in the width and height direction, and respectively represent the width and height of the real box; In the training process, the parameters of the model are updated by minimizing the SIoU loss function.
8. The method of claim 7, wherein: In step S42, the overall architecture of YOLOv8 includes three main parts: Backbone, Neck and Head; Backbone is responsible for extracting features from input images, Neck fuses and enhances features, and Head is used to predict the category and position of the target.
9. The method of claim 8, wherein: The CBAM attention mechanism module includes a channel attention module and a spatial attention module.
10. The method of claim 9, wherein: The channel attention module assigns different weights to each channel by operating on the channel dimension of the feature map to highlight important channel features; first, the input feature map respectively, global average pooling and global maximum pooling are performed to obtain two global feature vectors and The calculation formula is as follows: ; ; wherein, represents a feature map F values of all channels at the i-th row, j-th column; is a height of the feature map in the spatial dimension; is a width of the feature map in the spatial dimension; The two global feature vectors are respectively input into a shared multi-layer perception, and after ReLU activation function and Sigmoid activation function, a channel attention map M is obtained c ∈R C : ; wherein, σ represents a Sigmoid activation function; v avg is a global feature vector obtained by performing global average pooling on the input feature map F; v max is a global feature vector obtained by performing global maximum pooling (GMP) on the input feature map F; Finally, the channel attention map M c is multiplied with the input feature map F channel by channel to obtain a feature map enhanced by channel attention : ; wherein M c With F Element-wise multiplication.
11. The method of claim 10, wherein: The spatial attention module then operates on the spatial dimension of the feature map, assigns different weights to each spatial position, and highlights important spatial regions; first, the feature map enhanced by channel attention The channel dimension is respectively averaged and maximally pooled to obtain two spatial feature maps and The calculation formula is as follows: ; ; The two spatial feature maps are spliced in the channel dimension to obtain a new feature map , and a spatial attention map is obtained through a 7x7 convolution layer and a sigmoid activation function : ; wherein, represents a concatenation operation in the channel dimension; Finally, the spatial attention map is multiplied element-wise with the channel attention enhanced feature map to obtain the final CBAM attention mechanism enhanced feature map : ; wherein with element-wise multiplication.
Citation Information
Patent Citations
AI-based industrial image detection device and method
CN110602462A
Bearing surface scratch detection method based on machine vision
CN115272204A
Aircraft engine damage detection method and device based on artificial intelligence algorithm
CN118115909A
Overhead transmission line defect detection method based on multi-scale feature fusion
CN119624922A
Intelligent borescope detection system for aero-engine
CN120064292A