Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

237 results about "Information capture" patented technology

Information capture is the process of collecting paper documents, forms and e-documents, transforming them into accurate, retrievable, digital information, and delivering the information into business applications and databases for immediate action.

Flexible job shop scheduling method based on graph neural network and deep reinforcement learning

The invention discloses a flexible job shop scheduling method based on a graph neural network and deep reinforcement learning, and relates to the technical field of flexible job shop scheduling. The method at least comprises the following steps: S1, firstly, carrying out Markov Decision Process (MDP) on a flexible job shop scheduling problem, namely, FJSP, and initializing a scheduling state; and S2, representing a complex relationship between a job and a machine by using a heterogeneity graph, and effectively mapping different entities (the job, the machine, the operation and the like) of the problem and the relationship between the different entities into a graph structure, wherein the different entities (the job, the machine, the operation and the like) of the problem and the relationship between the different entities (the job, the machine, the operation and the like) of the problem are represented by the heterogeneity graph. According to the method, the graph neural network based on the meta-relationships is provided, different graph convolution modes are innovatively adopted for different meta-relationships to extract features, original semantic information is reserved, the global information capturing capability is enhanced, and a reinforcement learning agent is more accurate when making a scheduling decision.
Owner:CHONGQING UNIV OF TECH

Service path tracking and fault tracing method and device for video monitoring system

The invention provides a business path tracking and fault tracing method and device for a video monitoring system, and relates to the technical field of video monitoring fault diagnos.The video monitoring system is divided into a plurality of layers, and multiple types of data acquisition probes are deployed at key nodes of each layer; when a user initiates a video request, generating a Trace ID to be spread along with a video streaming signaling, and reconstructing an end-to-end service path in combination with node information captured by a probe; associating and fusing the multi-dimensional performance data by taking the TraceID as a key, and establishing an end-to-end total delay decomposition and dynamic health baseline; and positioning a root cause node based on the weighted directed graph, the node anomaly confidence score and the fault propagation consistency, displaying a path, data and a traceability result through a visual interface, and generating a precise alarm. According to the method, service link visualization is realized, traditional equipment alarm is upgraded to service fault root cause positioning, the troubleshooting efficiency is improved, and the method is suitable for an isomerized large-scale video monitoring system.
Owner:FNETLINK SMART CORE (HANGZHOU) TECHNOLOGY CO LTD

Language-guided structure-aware network architecture and method for camouflage target detection

The invention discloses a language-guided structure-aware network architecture and method for camouflage target detection, and relates to the technical field of camouflage target detection. According to the method, the text is guided to focus on a potential target, text semantic and visual features are fused through the CLIP model, the target mask is generated, the problems that a traditional model lacks semantic guidance and is difficult to focus on a camouflage area are solved, and background interference is greatly reduced; edge details can be accurately extracted, a Fourier edge enhancement module (FEEM) is combined with spatial domain edge enhancement and frequency domain high-frequency information capture, the problem of a fuzzy pain point of a camouflage target boundary is effectively solved, and the edge positioning precision is improved; structural and local dual optimization is proposed, a structural awareness attention module (SAAM) is fused with semantic and edge information, and a coarse guidance local refinement module (CGLRM) is designed through double branches of global guidance and local refinement, so that the regional consistency and structural integrity of a segmentation result are guaranteed, and local details are prevented from being lost.
Owner:CHONGQING UNIV OF TECH

Multi-mode weak supervision video anomaly detection method and device, equipment and medium

The invention belongs to the field of computer science and technology, and particularly relates to a multi-mode weak supervision video anomaly detection method and device, equipment and a medium. The method comprises the following steps: firstly, carrying out dynamic modeling on inter-fragment information by utilizing an external attention mechanism to obtain cross-fragment global context information; the temporal context aggregation module and the multi-scale time network are used to capture global and local information of the visual information and the text information within the fragment, and a feature representation containing local context details and the global information is generated. In addition, a multi-modal adaptive fusion mode is adopted, key modal features are focused in combination with target weights, further processing is performed through a multi-scale convolution attention module, and feature representation with higher discrimination is extracted. According to the method, the dependence on accurate annotation data in traditional video anomaly detection is effectively reduced; through hierarchical context modeling and an adaptive attention mechanism, the feature expression ability and the key information capture efficiency are enhanced, and a reliable anomaly detection solution is provided for an intelligent monitoring system.
Owner:TAIYUAN UNIVERSITY OF SCIENCE AND TECHNOLOGY

Foundation pit early warning method based on multi-source monitoring data and dynamic evaluation

The invention discloses a foundation pit early warning method based on multi-source monitoring data and dynamic evaluation, which relates to the technical field of foundation pit safety early warning and comprises the steps of S1, multi-element monitoring, S2, data processing and normalization, S3, dynamic safety level establishment, S4, dynamic weight and credibility calculation, S5, evidence fusion and safety level judgment and S6, dynamic risk evaluation and calibration. According to the method, all-dimensional data acquisition and high-quality data processing are formed through the steps S1 and S2, all-dimensional information capture of foundation pit engineering characteristics, geological conditions, surrounding environments and body states is achieved, data quality is guaranteed through data cleaning, normalization and validity verification, and the problems of monitoring data fragmentation and uneven quality are avoided; through the steps S3 and S5, a security level judgment system of dynamic probability allocation and multi-source evidence fusion is constructed, limitation of a static threshold method is prevented, and scientificity and accuracy of security level judgment are improved.
Owner:ZHENJIANG SURVEYING & MAPPING RES INST CO LTD

Product style trend analysis method and system driven by multi-modal neural network

The invention discloses a multi-modal neural network-driven product style trend analysis method and system, and relates to the technical field of image processing, and the method comprises the steps: receiving product multi-modal data, including text, image and audio data, carrying out the tensor extraction of the product multi-modal data, enabling each modal data to correspond to a single-peak tensor, and carrying out the tensor extraction of the product multi-modal data; carrying out feature fusion on the single-peak tensor to obtain a multi-modal feature tensor, and then carrying out dimension reduction processing; calculating a cross-modal multi-head attention score by using the single-modal tensor based on a language tensor, wherein the cross-modal multi-head attention score is used for weighting single-modal features; fusing weighted features, dimensionality reduction tensors and manually labeled style labels, constructing a supervised training target, and generating single-mode pseudo labels for multi-task model training; cross-modal shared information capture is carried out on the unimodal tensor based on a covariance matrix, cross-modal shared information is constrained through modal alignment loss, and under a multi-task learning framework, a trained multi-task model takes output of a multi-modal main task as a final style label to realize product trend analysis.
Owner:SUZHOU UNIV

Text-video retrieval method and system based on multi-granularity learnable interaction

The invention discloses a text-video retrieval method and system based on multi-granularity learnable interaction, and relates to the field of data retrieval, and the method comprises the following steps: encoding a query text and an unpaired video in to-be-retrieved data, and extracting query text features and video features; obtaining a plurality of learnable vectors, carrying out cross-modal interaction on the learnable vectors and fine-grained video features, carrying out cross-grained alignment on query text features and learnable variables, carrying out fine-grained alignment on each word feature and the learnable vector, and aligning the query text features and video features; and capturing fine-granularity alignment information and coarse-granularity alignment information, capturing weights of different similar vectors among modals, summing multiple similarity scores based on the weights to obtain final multi-granularity similarity scores, and sorting the final multi-granularity similarity scores to obtain a final retrieval result. According to the method, cross-modal interaction and other operations are carried out through the plurality of learnable vectors and the fine-grained visual features, the fine-grained visual information is utilized, and the retrieval accuracy is improved.
Owner:NINGXIA UNIVERSITY

Systems and methods for capturing and generating panoramic three-dimensional models and images

An environmental capture system (ECS) captures image data and depth information in a 360-degree scene. The captured image data and depth information can be used to generate a 360-degree scene. The ECS comprises a frame, a drive train mounted to the frame, and an image capture device coupled to the drive train to capture, while pointed in a first direction, a plurality of images at different exposures in a first field of view (FOV) of the 360-degree scene. The ECS further comprises a depth information capture device coupled to the drive train. The depth information capture device and the image capture device are rotated by the drive train about a first, substantially vertical, axis from the first direction to a second direction. The depth information capture device, while being rotated from the first direction to the second direction, captures depth information for a first portion of the 360-degree scene. The image capture device captures, while pointed in the second direction, a plurality of images at different exposures in a second FOV that overlaps the first FOV of the 360-degree scene. The depth information capture device and the image capture device are rotated by the drive train about the first axis from the second direction to a third direction. The depth information capture device, while being rotated from the second direction to the third direction, captures depth information for a second portion of the 360-degree scene. The image capture device, while pointed in the third direction, captures a plurality of images at different exposures in a third FOV that overlaps the second FOV of the 360-degree scene.
Owner:COSTAR REALTY INFORMATION INC

A multi-stage adjustment mechanism for a prestressed positioning mesh

The present application relates to the technical fields of prestressed positioning mesh welding, and especially relates to a multi-stage adjusting mechanism for prestressed positioning mesh, which is provided with a welding control module, an information capturing module, a displacement analysis module and a welding correction module. The transverse steel bars are laid on the longitudinal steel bars through the laying unit, the intersection positions of the longitudinal steel bars and the transverse steel bars are welded through the welding unit, the vibration intensity change curve is drawn through the information capturing module, it is determined whether there is vibration allowance abnormal phenomenon at the intersection positions and the vibration allowance abnormal category of the intersection positions is determined through the abnormal analysis unit, the welding correction mode is selected through the welding correction module, and then, the quick abnormal risk identification according to the actual situation in the steel bar welding and sticking process is realized, the abnormal welding positions are classified and processed, the welding correction mode is adaptively adjusted, and the production efficiency and welding reliability of the prestressed positioning mesh are improved.
Owner:CHINA TIESIJU CIVIL ENGINEERING GROUP CO LTD +1

Systems and methods for private authentication with helper networks

Helper neural network can play a role in augmenting authentication services that are based on neural network architectures. For example, helper networks are configured to operate as a gateway on identification information used to identify users, enroll users, and / or construct authentication models (e.g., embedding and / or prediction networks). Assuming, that both good and bad identification information samples are taken as part of identification information capture, the helper networks operate to filter out bad identification information prior to training, which prevents, for example, identification information that is valid but poorly captured from impacting identification, training, and / or prediction using various neural networks. Additionally, helper networks can also identify and prevent presentation attacks or submission of spoofed identification information as part of processing and / or validation.
Owner:PRIVATE IDENTITY LLC

Systems and methods for capturing and generating panoramic three-dimensional models and images

An environmental capture system (ECS) captures image data and depth information in a 360-degree scene. The captured image data and depth information can be used to generate a 360-degree scene. The ECS comprises a frame, a drive train mounted to the frame, and an image capture device coupled to the drive train to capture, while pointed in a first direction, a plurality of images at different exposures in a first field of view (FOV) of the 360-degree scene. The ECS further comprises a depth information capture device coupled to the drive train. The depth information capture device and the image capture device are rotated by the drive train about a first, substantially vertical, axis from the first direction to a second direction. The depth information capture device, while being rotated from the first direction to the second direction, captures depth information for a first portion of the 360-degree scene. The image capture device captures, while pointed in the second direction, a plurality of images at different exposures in a second FOV that overlaps the first FOV of the 360-degree scene. The depth information capture device and the image capture device are rotated by the drive train about the first axis from the second direction to a third direction. The depth information capture device, while being rotated from the second direction to the third direction, captures depth information for a second portion of the 360-degree scene. The image capture device, while pointed in the third direction, captures a plurality of images at different exposures in a third FOV that overlaps the second FOV of the 360-degree scene.
Owner:COSTAR REALTY INFORMATION INC

Occupancy map segmentation for autonomous guided platform with deep learning

The technology disclosed includes systems and methods for preparing a segmented occupancy grid map based upon image information of an environment in which a robot moves. The image information is captured by at least one visual spectrum-capable camera and at least one depth measuring camera. The system includes logic to receive image information captured by at least one visual spectrum-capable camera and location information captured by at least one depth measuring camera located on a mobile platform. The system includes logic to extract from the image information, features in the environment. The system includes logic to determine a 3D point cloud of points having 3D information. The system includes logic to determine, from the 3D point cloud, an occupancy map of the environment. The system includes logic to segment the occupancy map into a segmented occupancy map of regions that represent rooms and corridors in the environment.
Owner:TRIFO INC

Target tracking early warning method and system based on radar photoelectric linkage and information fusion

The invention discloses a target tracking early warning method and system based on radar photoelectric linkage and information fusion, and relates to the technical field of electric digital data processing, and the method comprises the steps: obtaining the information of a moving target, and obtaining a real-time azimuth angle and a distance sequence; obtaining an absolute difference value between adjacent distance data in the distance sequence, and obtaining a stability index; generating a correction value, obtaining a compensation azimuth angle range, obtaining a steering angle range of the photoelectric equipment, obtaining a compensation distance range, and obtaining optical zoom information of the photoelectric equipment; and driving photoelectric equipment to capture target contour features according to the steering angle range and the optical zoom information to lock a target, and calling a physical rule base to obtain a target classification and early warning strategy based on the moving target information obtained by the radar and the target contour features captured by the photoelectric equipment. The method has the advantages of dynamic self-adaption, high resource utilization rate and cooperative closed loop.
Owner:XIAN CHENHANG EXCELLENCE TECH CO LTD

Crack detection method based on multi-scale low-rank fusion attention network

The invention discloses a crack detection method based on a multi-scale low-rank fusion attention network. The method comprises the following steps: S1, acquiring a bridge crack image through image acquisition equipment, and constructing a bridge crack data set with labels; s2, constructing a low-rank self-attention fusion module LSAFS3, constructing a dilated convolution attention fusion module DC-AFM, performing a synergistic effect with the low-rank self-attention fusion module LSAF, capturing multi-scale features through dilated convolution with different expansion rates, realizing feature optimization fusion in combination with an attention mechanism, and constructing a multi-scale low-rank fusion attention network model; s4, based on the data set constructed in S1, training and optimizing the multi-scale low-rank fusion attention network model jointly constructed in S2 and S3 to obtain a final bridge crack detection model; detecting a to-be-detected bridge crack image by using the model, and outputting a detection result; the method has the characteristics of capturing crack information in a multi-scale manner and improving the precision and efficiency of crack detection.
Owner:XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY

Object scale utilizing away-facing images

ActiveUS12670649B2Stereo cameraRadiology
Methods, systems, mobile devices, and non-transitory computer-readable mediums for determining a scale of an object with a mobile device. The mobile device includes at least one object-facing camera (e.g., a monocular camera) and an away-facing stereo camera system. Information captured using the away-facing stereo camera system is used to estimate relative pose information for determining the scale of the object in images captured by the object-facing camera. The proposed system leverages the scene images surrounding the mobile device to resolve scale ambiguity and calculate the absolute scale.
Owner:SNAP INC

Multimodal motion vision based authentication system a method thereof

The present invention discloses a motion vision-based multimodal authentication system and a method thereof. The system comprises several components, including an information-capturing module, an AI module, a memory unit, a processor unit, and a temporal authentication module. The system captures information data from the user and the user's surroundings, categorize the data into identity and predetermined categories, generates login credentials or authorization keys, and matches the real-time information data for user authorization.
Owner:HUMMINGBIRDS AI INC

An image processing method and system for embankment inspection

ActiveCN121438152BEnsure inspection efficiencyquality improvementImaging processingMirror reflection
The present application relates to the field of dam inspection, especially to a kind of image processing method and system for dam inspection, the present application is by controlling unmanned aerial vehicle to fly over detection area according to predetermined trajectory, when unmanned aerial vehicle is pre-entered detection area, continuously obtain image information and carry out multidimensional analysis, based on the result of multidimensional analysis and mirror surface reflection index, calculate to obtain spectral characteristic interference representation value, set interference label for detection area, subsequent interference label based on the control of unmanned aerial vehicle in detection area adaptively obtains spectral information in monitoring area, subsequent generation perspective adjustment instruction, by sampling determination selection confidence sampling position obtains spectral information, the present application is by image information in advance sensing interference situation of spectral information extraction in detection area, adaptively select spectral information capture mode, under the premise of guaranteeing inspection efficiency, improve the quality and reliability of the obtained spectral information, provide effective data support for subsequent analysis dam potential anomaly.
Owner:WATER RESOURCES RES INST OF SHANDONG PROVINCE

Interactive narrative method and device based on artificial intelligence and storage medium

The invention provides an interactive narrative method and device based on artificial intelligence, and a storage medium. The method comprises the following steps: receiving input information from a user; capturing a current user interaction event according to the input information; determining a transfer condition set corresponding to the current plot state based on a finite plot state machine; determining a state transition condition satisfied by the current user interaction event from a transition condition set, and transitioning from the current plot state to a target plot state corresponding to the state transition condition; and executing the plot logic corresponding to the target plot state to present the plot content to the user. Therefore, the narrative process is constrained in the structured narrative graph defined by the finite plot state machine, thereby realizing effective guidance and control of the narrative macroscopic trend while maintaining the interactive openness, ensuring the continuity and logicality of the narrative, and improving the success rate of the narrative process. The problems of uncontrollable plot development and poor logic reliability in the interactive narrative process in related technologies are effectively solved.
Owner:HUACUI STARLIGHT (BEIJING) INTELLIGENT TECHNOLOGY CO LTD

Industrial part surface defect detection method based on multi-mode and self-adaption

The invention discloses a multi-modal and adaptive industrial part surface defect detection method, which comprises the following steps of: inputting image information shot on the surface of an industrial part into a feature extraction network of an FPN (Fabry-Perot Network) structure to obtain initial feature maps of different hierarchies from bottom to top, and fusing the initial feature maps from top to bottom to obtain fused feature maps of different hierarchies; a reverse fusion path is added in a feature extraction network to obtain reverse fusion feature maps of different hierarchies, and the method has the advantages that the feature extraction network of an existing FPN structure is improved, the reverse fusion path is added, and the fusion feature maps are reversely fused from bottom to top to obtain the reverse fusion feature maps of different hierarchies; the accuracy of industrial part surface defect detection is greatly improved, and the false detection and omission ratio is reduced.
Owner:ZHEJIANG WANLI UNIV

Complex environment three-dimensional scene reconstruction method based on multi-view neural network

The patent application provides a complex environment three-dimensional scene reconstruction method based on a multi-view neural network, and the method mainly comprises two steps: scene information capturing: obtaining scene data from a plurality of views through a high-resolution camera, and ensuring that all details of a complex scene are covered; the captured image is standardized, including color correction and noise removal, to ensure image quality and consistency. And three-dimensional model reconstruction: performing three-dimensional reconstruction on the acquired scene data by using a neural network. Firstly, light sampling is performed on a scene through camera parameters and pose information, and an occupancy grid is created to optimize the sampling efficiency. Thirdly, decomposing the sampling data into feature vectors and inputting the feature vectors into a neural network which is divided into a density network and a color network and is used for generating volume density and color information of the scene; and finally, obtaining a three-dimensional scene model through a volume rendering technology. The method has the advantage that a high-precision three-dimensional model can be quickly generated in a complex environment through an efficient sampling and data processing technology.
Owner:XINJIANG UNIVERSITY

A mount security detection method

A mount security detection method for an in-vehicle information capture device having at least one acceleration sensor to gather acceleration data, and comparing a sample of the acceleration data for
Owner:APPY RISK TECH LTD

Conference recording apparatus, conference recording method, and conference recording program

To make it easy to perform settings for suppressing noise or the like contained in collected voice.SOLUTION: A conference recording apparatus receives, from a conference device equipped with a microphone that collects voice arriving from the surroundings and a camera that captures images of the surroundings, collected sound information collected by the microphone and video information captured by the camera, and displays a first setting area corresponding to directions of the surroundings together with the video information. A first range is set for the first setting area, and voice components arriving from directions within the first range in the collected sound information are suppressed.SELECTED DRAWING: Figure 3
Owner:PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD

Enhancing privacy in biometrics-based authentication systems

According to one embodiment, a method, computer system, and computer program product for privacy-enhanced, biometrics-based authentication is provided. The embodiment may include capturing credential information on a credential medium and biometric information for a user, each provided by the user, by one or more information capture devices. The embodiment may also include identifying a hash stored on the credential medium based on the captured credential information. The embodiment may further include calculating, locally, a hash of the biometric information using a preconfigured hashing algorithm. The embodiment may also include, in response to the identified hash matching the calculated hash, authenticating the user.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

Insulator defect detection method and device

The invention provides an insulator defect detection method and device, and belongs to the field of image recognition. The method comprises the following steps: acquiring a to-be-detected image of a target insulator; inputting the to-be-detected image into a pre-trained detection model; wherein the pre-trained detection model comprises an input end, a backbone network, a neck network and a prediction network; the backbone network comprises a multi-scale channel information extraction module and a global-local attention module; the input end is used for inputting a to-be-detected image into the backbone network, and the backbone network is used for extracting feature information of the to-be-detected image to obtain a feature map; the neck network is used for fusing the feature maps to obtain a fused feature map; the prediction network is used for performing prediction based on the fused feature map to obtain a predicted image; and determining a detection result of the target insulator based on the prediction image. The feature information capturing capability and the detection precision of the insulator defect detection method based on the YOLO series algorithm can be improved.
Owner:XINGTAI LIGE ELECTRICAL TECHNOLOGY CO LTD

A method and system for predicting quality indicators of a catalytic cracking process

The application provides a catalytic cracking process quality index prediction method and system, and the method comprises the following steps: obtaining actual production data of a catalytic cracking process in a preset time, pre-processing and selecting features of the data, constructing an input data set of a convolutional neural network for prediction, and confirming a prediction index; inputting the data set into a pre-established convolutional neural network, firstly extracting input feature information by using one-dimensional convolution, retaining the feature information most related to the prediction value in the convolution layer by using a maximum pooling technology, taking a bidirectional gated recurrent unit as a basic unit, processing the output information of the pooling layer in a time forward and backward direction, capturing the time dependence between the features, and selecting the key information extracted from the convolution layer and the pooling layer by using an attention layer, and then inputting the key information into a full connection layer to obtain a prediction value. Based on the method, a catalytic cracking process quality index prediction system is further provided. The application greatly improves the prediction accuracy of the catalytic cracking process quality index.
Owner:CHINA PETROLEUM & CHEMICAL CORP +1

Camera module and intelligent electronic equipment

PendingCN121334472AEngineeringCamera module
The invention discloses a camera module and intelligent electronic equipment. Wherein the camera module comprises a first air bag and a camera module, the first air bag can be fixed to the head-mounted device, the first air bag is provided with a first side wall, the wall thickness of the first side wall is gradually reduced in the first direction, and the first air bag has an initial state and an inflation state; the camera module is arranged on the first side wall and is used for acquiring image information and transmitting the image information to the head-mounted device; when the first air bag is switched between the initial state and the inflation state, the angle of the camera module relative to the head-mounted device can be adjusted. When the camera module provided by the invention is applied to the head-mounted device, the camera module has a larger visual range and image information capturing accuracy.
Owner:BEIJING GOERTEK TECH CO LTD

Target tracking and early warning method and system based on radar photoelectric linkage and information fusion

This invention discloses a target tracking and early warning method and system based on radar-optical-electric linkage and information fusion, belonging to the field of electronic digital data processing technology. The method includes: acquiring moving target information, acquiring real-time azimuth angle and range sequence; acquiring the absolute difference between adjacent distance data in the range sequence, acquiring stability indicators; generating correction values, acquiring the compensation azimuth angle range, acquiring the turning angle range of the optoelectronic device, acquiring the compensation distance range, acquiring the optical zoom information of the optoelectronic device; driving the optoelectronic device to capture target contour features and lock the target according to the said turning angle range and the said optical zoom information; and based on the moving target information acquired by radar and the target contour features captured by the optoelectronic device, calling a physical rule base to obtain target classification and early warning strategies. This invention has the advantages of dynamic adaptation, high resource utilization, and collaborative closed loop.
Owner:XIAN CHENHANG EXCELLENCE TECH CO LTD

Secure data processing using data packages generated by edge devices

Disclosed are example methods, systems, and devices that allow for secure data processing using data packages generated by edge devices. The techniques include generating a biometric signature using information captured by a computing device, and encrypting user data obtained via the computing device using the biometric signature and a device identifier of the computing device. A security token can be generated and utilized by the computing device to generate a data package, which is configured such that any change to the data package would cause a validation process of the data package using the security token to fail. The data package can be encrypted using various digital keys and provided to secondary computing systems.
Owner:WELLS FARGO BANK NA

A photovoltaic panel defect category detection method based on feature pyramid and cascaded group attention

The present application relates to the field of computer vision, and discloses a photovoltaic panel defect category detection method based on feature pyramid and cascade group attention, comprising the following steps: S1. Data acquisition, labeling and preprocessing; S2. Model training using improved YOLOv11 algorithm; S3. Predicting the input infrared picture, accurately positioning the defect and labeling the category. The photovoltaic panel defect category detection algorithm based on feature pyramid and cascade group attention according to the internal circuit structure of the photovoltaic panel and the infrared imaging picture accurately positions, labels and classifies the defects on the photovoltaic panel; the latest YOLOv11 target detection algorithm is adopted and improved, part of the C3k2 module is replaced with SCC3k2, so that the information capture on space and channel is more accurate, and the adverse effects caused by redundant information extraction are reduced; at the same time, different input segmentation is provided for each attention head in the attention module, and cross-head cascade output features are used to enhance the feature diversity of the input to the attention head.
Owner:GUANGDONG UNIV OF TECH