Marine regenerated PET plastic recovery treatment system and treatment process

Through multi-agent system and multi-spectral image enhancement technology that coordinates sea, land and air, the problem of low recognition and recycling of marine PET plastics is solved, and efficient and accurate identification and recycling in complex marine environments is achieved.

CN120297961AInactive Publication Date: 2025-07-11YILIAN PLASTICS SHENZHEN CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510486172.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot efficiently and accurately identify and recycle PET plastics in marine environments, especially in complex and multi-scenarios with low recognition accuracy and poor environmental adaptability, and traditional systems are difficult to cope with the impact of underwater shading and pollution.

Method used

Build a multi-agent system with sea, land and air coordination, combines underwater multi-spectral image enhancement network, self-attention detection architecture and 3D voxelization representation, and realizes accurate identification and recycling of PET plastics through a multi-task learning framework, adopts a self-organized map model to predict the distribution density, and uses a reinforced learning multi-agent collaborative grabbing control system for efficient recycling.

Benefits of technology

Active source search, accurate identification and efficient recycling of marine PET plastics has been achieved, with an identification accuracy of 92.2%, and a 3.5-fold increase in recycling efficiency, adapting to multi-scene recognition and recycling in complex marine environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120297961A_ABST
    Figure CN120297961A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of plastic recovery, and discloses a marine regenerated PET plastic recovery processing system and a processing technology.The marine regenerated PET plastic recovery processing technology comprises the steps that a multi-agent system is constructed through a water surface unmanned ship, an underwater robot and an air unmanned aerial vehicle, collaborative decision making and task allocation are conducted based on a behavior tree framework, and a target object is obtained; a self-organizing map model is adopted to predict the PET plastic distribution density, and a multi-agent system is guided to perform efficient source searching; according to the invention, active source searching, accurate identification and efficient recovery of PET plastics in a marine environment are realized by constructing a sea-land-air cooperative multi-agent system, an underwater PET plastic multi-target detection model and a stained PET plastic feature extraction and classification system, so that the technical problems of low identification accuracy and low recovery efficiency in a traditional recovery technology are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of plastic recycling and treatment, and more specifically, it relates to a marine recycled PET plastic recycling and treatment system and treatment process. Background Art

[0002] The formation of marine PET plastic pollution is mainly due to the combined effect of the following multi-dimensional reasons: synthetic textile wastewater discharge: polyester (PET), as the main component of synthetic textiles, releases fibers during the washing process, which enter the ocean through the wastewater system. Up to 94.8% of the microplastics are in the form of fibers; lack of plastic product management: difficult-to-recycle PET products such as agricultural films and packaging bags enter the ocean through river scouring, becoming a long-term pollution source; in addition, the phenomenon that landfill sites on land are washed into the ocean during storms also exacerbates the pollution; improper disposal in the shipping industry: plastic waste generated during shipping (such as packaging materials, discarded fishing gear, etc.) is directly dumped. According to statistics, 70% of marine plastic waste comes from such behaviors; marine accidents: when a cargo ship is in distress, containers loaded with PET products may fall into the sea, resulting in a large amount of plastic entering the marine environment at once.

[0003] Marine PET plastic pollution has become a global environmental problem, and the existing technologies have the following deficiencies: traditional recycling systems rely on passive monitoring and lack the ability to actively search for sources; single-environment deployment is difficult to cope with complex situations distributed in multiple scenarios such as the sea surface, underwater, and coastal areas; multi-spectral recognition systems need to identify materials after they are taken out of the water and cannot directly identify various PET plastics in the underwater environment in real time; when multiple types of plastics are stacked or partially blocked, the recognition accuracy drops significantly; after being soaked in the ocean, the surface of PET plastics is attached with biofilms, salts, and pollutants, resulting in low accuracy of traditional optical recognition systems and inability to effectively distinguish PET plastics with different degradation degrees. Summary of the Invention

[0004] The present invention provides a marine recycled PET plastic recycling and treatment system and treatment process to solve the technical problems of marine recycled PET plastic recycling and treatment in related technologies.

[0005] The present invention provides a marine recycled PET plastic recycling and treatment process, including the following steps:

[0006] Construct a multi-agent system through surface unmanned boats, underwater robots, and aerial drones, conduct collaborative decision-making and task allocation based on the behavior tree framework, and use the self-organizing map model to predict the distribution density of PET plastics to guide the multi-agent system to conduct efficient source search;

[0007] Utilize a multi-spectral image enhancement network and self-attention detection architecture for the underwater environment, combined with binocular stereo vision to construct a 3D voxel representation to achieve precise identification of underwater PET plastics;

[0008] Collect plastic feature data through multispectral imaging technology, and combine a multi-task learning framework to simultaneously identify PET material types, evaluate degradation degrees, and perform morphological analysis;

[0009] The multi-agent collaborative grasping control system based on reinforcement learning realizes the precise recycling of PET plastics.

[0010] As a further optimization scheme of the present invention, the multi-agent system adopts a cross-media perception information fusion algorithm to fuse multi-source data from the air, water surface, and underwater. The fusion process is expressed as:

[0011] D fused = f fusion (D air , D surface , D underwater );

[0012] Among them, D fused represents the fused perception data, D air , D surface and D underwater represent the air, water surface, and underwater perception data respectively, and f fusion is the fusion function.

[0013] As a further optimization scheme of the present invention, the self-organizing map model calculates the PET plastic distribution density through the following steps:

[0014]

[0015] Among them, ρ(x, y, z, t) is the predicted value of the PET plastic density at the spatial point (x, y, z) at time t, represents the weighted sum of k basis functions, and φ i (x, y, z, t) represents the i-th basis function, which is used to describe the space of plastic distribution, is the corresponding weight coefficient, and k is the total number of basis functions. The weight coefficient is continuously updated through online learning to adapt to the changing marine environment.

[0016] As a further optimization scheme of the present invention, the multi-spectral image enhancement network adopts a generative adversarial network architecture, including:

[0017] G network (I raw ) = I enhanced ;

[0018] D network (I) ∈ [0, 1];

[0019] Among them, G networkis the generator network, I raw is the original underwater image, I enhanced is the enhanced image, D network is the discriminator network, which outputs the authenticity score of the image, and I is the input image.

[0020] As a further optimization scheme of the present invention, the 3D voxelization representation obtains depth information through binocular stereo vision to construct a voxel grid V:

[0021]

[0022] where (x, y, z) represents a point in three-dimensional space.

[0023] As a further optimization scheme of the present invention, the multi-task learning framework simultaneously optimizes the loss functions of three tasks:

[0024]

[0025] where and respectively represent the loss functions of material type recognition, degradation degree evaluation, and morphological analysis, and λ1, λ2, and λ3 are the weight coefficients of each task.

[0026] As a further optimization scheme of the present invention, the multi-agent collaborative grasping control system of reinforcement learning optimizes the policy by maximizing the cumulative reward function:

[0027]

[0028] where π represents the policy function, represents the cumulative summation from time step 0 to T RL E represents the expected value operator, argmax represents the variable value corresponding to the maximum value of the objective function, and π * is the optimal policy, r t is the reward at time step t, represents the discount factor at time t, and T RL is the task completion time. This optimization process ensures the collaborative efficiency of the multi-agent system in complex marine environments.

[0029] As a further optimization scheme of the present invention, the multi-scale feature fusion mechanism is calculated according to the following formula:

[0030]

[0031] where represents the summation operation from i = 1 to n, where i is the summation variable and n is the upper limit of the summation, and F i is the feature map of the i-th scale, and Wi is the corresponding transformation matrix, α i is the adaptive weight coefficient, F fused is the fused feature map.

[0032] As a further optimized solution of the present invention, it further includes a temporal attention mechanism step for capturing the motion characteristics of plastics in water flow:

[0033] A t = softmax(f query (F t )·f key (F t-k:t-1 ) T )·f value (F t-k:t-1 );

[0034] Among them, softmax represents the softmax function for converting a vector into a probability distribution, and A t is the attention feature at time step t, F t is the feature at the current time step, F t-k:t-1 is the feature of the previous k time steps, and f query , f key and f value are the query, key, and value mapping functions respectively, and T represents the matrix transpose operator.

[0035] The present invention provides a marine recycled PET plastic recycling and processing system for implementing the above-mentioned marine recycled PET plastic recycling and processing process, including:

[0036] A multispectral imaging module for collecting multi-band spectral feature data of underwater PET plastics;

[0037] A GPS positioning module for obtaining the precise position information of each intelligent agent in real time;

[0038] A surface maneuvering execution module responsible for the motion control of the surface unmanned boat;

[0039] A binocular stereo vision module for constructing a 3D voxel representation of underwater targets;

[0040] An underwater propulsion and grasping module responsible for the motion and grasping actions of the underwater robot respectively;

[0041] A flight control module responsible for the flight path planning and attitude control of the aerial drone;

[0042] An image enhancement module for improving the quality of underwater images through a generative adversarial network;

[0043] A target detection module for accurately positioning PET plastics based on the self-attention mechanism;

[0044] The material analysis module uses a multi-task learning framework to evaluate plastic characteristics;

[0045] The multi-agent collaborative decision-making module performs task allocation based on the behavior tree framework;

[0046] The task planning module is responsible for formulating global search and recovery strategies;

[0047] The data processing module is used to perform fusion analysis on multi-source sensing data.

[0048] The beneficial effects of the present invention are as follows: By constructing a multi-agent system for sea-land-air collaboration, an underwater PET plastic multi-object detection model, and a system for feature extraction and classification of fouled PET plastics, the present invention realizes the active source seeking, precise identification, and efficient recovery of PET plastics in the marine environment; at the same time, uses the self-organizing map model to predict the plastic distribution density, improves the underwater recognition accuracy through the multi-spectral image enhancement network and 3D voxel representation, uses the multi-task learning framework to realize the synchronous analysis of plastic material types, degradation degrees, and morphologies, and realizes precise recovery through the multi-agent collaborative grasping control system based on reinforcement learning. In this way, it solves the technical problems such as low recognition accuracy, poor environmental adaptability, and low recovery efficiency in traditional recovery technologies. The plastic recognition accuracy rate in complex marine environments reaches 92.2%, and the recovery efficiency is increased by 3.5 times, having significant technical effects and application values. Brief Description of the Drawings

[0049] Figure 1 is a flowchart of a marine recycled PET plastic recovery and treatment process of the present invention;

[0050] Figure 2 is a flowchart of the efficient source seeking of the multi-agent system of the present invention;

[0051] Figure 3 is a flowchart of the precise identification of underwater PET plastics of the present invention;

[0052] Figure 4 is a flowchart of the feature extraction and classification of fouled characteristics in the marine environment of the present invention;

[0053] Figure 5 is a flowchart of the precise recovery of PET plastics of the present invention. Detailed Embodiments

[0054] Reference will now be made to exemplary embodiments to discuss the subject matter described herein. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0055] Embodiment 1

[0056] The present invention discloses a recycling process for marine recycled PET plastics, as Figures 1 - 5 shown, which includes the following steps:

[0057] Step 1: Construct a multi-agent system through surface unmanned boats, underwater robots, and aerial drones, conduct collaborative decision-making and task allocation based on the behavior tree framework, and use the self-organizing map model to predict the distribution density of PET plastics to guide the multi-agent system to conduct efficient source search;

[0058] This step realizes the active source search and recycling of PET plastics across media and all terrains. Specifically, it includes the following sub-steps:

[0059] Sub-step 1.1: Construct a multi-agent collaborative decision-making framework based on the behavior tree;

[0060] This step constructs a multi-agent collaborative decision-making framework using a hierarchical behavior tree structure, processes input data from various unmanned systems (surface unmanned boats, underwater robots, and aerial drones), and outputs a collaborative task allocation plan for different environments. This framework includes three layers: Task planning layer: Responsible for global goal decomposition and task priority ranking; Role assignment layer: Conduct role matching based on the sensing capabilities and maneuverability of each agent; Action execution layer: Convert high-level instructions into specific motion control commands.

[0061] The collaborative decision-making process can be expressed as:

[0062] D(t) = f(S1(t), S2(t),..., S n (t), E(t));

[0063] where D(t) is the decision result at time t, S1(t), S2(t), S n (t) are the state information of the 1st, 2nd, nth agents respectively, E(t) is the environmental state, and f is the decision function, which integrates the conditional judgment and behavior selection logic of each node in the behavior tree.

[0064] Sub-step 1.2: Construct a cross-media perception information fusion algorithm;

[0065] Construct a perception information fusion algorithm for multi-source data in the air, on the water surface, and underwater, process multi-modal and heterogeneous sensing data, and output a unified environmental perception result. This algorithm includes: a spatio-temporal alignment module: to solve the problems of time and space inconsistency of different sensor data; a heterogeneous data conversion module: to convert different types of data into a unified feature space; a multi-level fusion module: to achieve fusion at three levels of data, features, and decisions.

[0066] The fusion process can be expressed as:

[0067] D fused = f fusion (D air , D surface , D underwater );

[0068] Among them, D fused represents the fused perception data, D air , D surface and D underwater represent the perception data in the air, on the water surface, and underwater respectively, and f fusion is the fusion function.

[0069] Sub-step 1.3: Construct a PET plastic distribution density prediction model based on self-organizing maps;

[0070] Construct a self-organizing map model using historical data and real-time observations, analyze the distribution law of marine PET plastics, and predict the most likely plastic aggregation areas. This model processes the fused multi-source sensing data and outputs a PET plastic distribution density prediction map in three-dimensional space.

[0071] This self-organizing map model consists of an input layer, a competitive layer, and an output layer. The input layer receives a vector containing features such as position coordinates (x, y, z), time t, current flow velocity, wind speed, and historical plastic density. The competitive layer contains k neurons, and each neuron corresponds to a weight vector w i , representing the spatial distribution pattern under specific environmental conditions. Through the iterative training process, the weight vector is continuously adjusted to adapt to the input data distribution:

[0072] w i (t + 1) = w i (t) + η(t) · h ci (t) · [x(t) - w i (t)];

[0073] Among them, η(t) is the learning rate, which gradually decreases as the training progresses; h ci (t) is the neighborhood function, indicating the influence of the winning neuron c on neuron i; x(t) is the current input vector; w i(t + 1) represents the weight vector of the i-th neuron at time t + 1; w i (t) represents the weight vector of the i-th neuron at time t.

[0074] In practical applications, this model can be used to predict the plastic aggregation points in the bay area after a typhoon. For example, in the Sanya Bay area, the system predicts that a plastic dense zone will appear in the reef area on the northeast side of the bay mouth by inputting the current wind direction (southeast wind, level 5), tide (ebb tide), and plastic point data found during the patrol in the past 24 hours. Based on this, the unmanned system preferentially arranges the patrol route for this area. The plastic aggregation area that takes 3 hours to discover in the conventional cruising mode can be accurately located only in 45 minutes through the prediction of this model, greatly improving the recovery efficiency.

[0075] The density prediction function is expressed as:

[0076]

[0077] Among them, ρ(x, y, z, t) is the predicted value of the PET plastic density at the spatial point (x, y, z) at time t, represents the weighted sum of k basis functions, φ i (x, y, z, t) represents the i-th basis function, which is used to describe the space of plastic distribution, is the corresponding weight coefficient, and k is the total number of basis functions. The weight coefficient is continuously updated through an online learning method to adapt to the changing marine environment.

[0078] Sub-step 1.4: Construct a multi-agent collaborative grasping control system based on reinforcement learning;

[0079] Construct a collaborative grasping control system that combines deep reinforcement learning and dynamic programming to process the real-time position and attitude data of the grasping target and output the collaborative grasping behavior control instructions for multiple agents. This system includes: a state representation module: encoding the grasping scene into a state vector; a Q-value network: evaluating the expected benefits of different grasping actions; a behavior selection module: selecting the optimal collaborative grasping strategy according to the current state.

[0080] This deep reinforcement learning model adopts a Double Deep Q-Network (DoubleDQN) structure, which consists of a target network and an evaluation network. Both have the same multi-layer structure: an input layer (state vector, including features such as the positions of each agent, the position and attitude of the plastic target), three fully connected layers (each layer contains 128 neurons, using the ReLU activation function), and an output layer (representing the Q-values of different collaborative grasping actions).

[0081] The update of network parameters follows the following rules:

[0082]

[0083] Among them, L RL (θ RL ) represents the reinforcement learning loss function to evaluate the network parameter θ RL is a variable, E represents the mathematical expectation, θ RL is the evaluation network parameter, Q(s,a) represents the value function of taking action a in state s, s′ represents the next state, and a′ represents the next possible action. is the target network parameter, r is the immediate reward, and γ RL is the discount factor, and argmax a′ represents selecting the action a′ that maximizes the Q value.

[0084] The target network parameter is softly updated using the evaluation network parameter every certain number of steps:

[0085]

[0086] Among them, τ is the update coefficient.

[0087] In practical applications, the system can effectively handle the collaborative recycling task of large marine PET plastic debris. For example, when faced with a group of PET plastics with an area of about 2.5 square meters and an irregular shape, the system will coordinate and command the surface unmanned boat and two underwater robots to carry out collaborative operations according to the real-time captured plastic positions, shapes, and floating states. The unmanned boat is responsible for blocking the drift path and stabilizing the target, while the underwater robots approach from different angles. One robot is responsible for fixing the plastic edge, and the other is responsible for capturing and curling the plastic main body. Tests show that compared with a single robot, the successful recovery rate of the collaborative system for large plastic debris has increased from 63% to 92%, and the average recovery time has been shortened by 57%.

[0088] The reinforcement learning model is optimized by maximizing the cumulative reward function:

[0089] Q(s,a) = R(s,a) + γ RL max a′ Q(s′,a′);

[0090] Among them, Q(s,a) is the value function of taking action a in state s, R(s,a) is the immediate reward, γ RL is the discount factor, max a′ represents selecting the action that maximizes the Q value among all possible next actions a′, and Q(s′,a′) represents the value function of taking action a′ in the next state s′.

[0091] Step 2: Use the multi-spectral image enhancement network and self-attention detection architecture for the underwater environment, combined with binocular stereo vision to construct a 3D voxel representation to achieve accurate identification of underwater PET plastics;

[0092] In this step, by constructing an image processing and target detection model for the special underwater environment, the problem of real-time recognition of PET plastics in complex environments such as underwater light scattering, turbid sea water, and object occlusion is solved. It specifically includes the following sub-steps:

[0093] Sub-step 2.1: Construct a multi-spectral image enhancement network for the underwater environment;

[0094] Construct a multi-spectral image enhancement network based on the generative adversarial network to process multi-spectral image data affected by scattering and absorption underwater and output a clear enhanced image. This network includes: a multi-scale encoder: extracting image features at different resolution levels; an attention-guided decoder: reconstructing the enhanced image and focusing on valuable regions; an adversarial discriminator: evaluating the authenticity of the generated image and promoting the improvement of the generation quality.

[0095] This multi-spectral image enhancement network uses an improved U-Net structure as the generator, and the discriminator uses the PatchGAN structure. The generator contains 5 downsampling layers and 5 upsampling layers. Each downsampling layer includes a convolutional layer (with a convolutional kernel size of 4×4 and a stride of 2), an instance normalization layer, and a LeakyReLU activation function. The upsampling layer includes a transposed convolution, an instance normalization, a Dropout (with a ratio of 0.5), and a ReLU activation function. A SkipConnection is added between the downsampling and upsampling to enhance feature transmission. In particular, the encoder part of the network can simultaneously process image inputs of different wavelengths (including the visible light band of 300 - 700nm and the near-infrared band of 700 - 2500nm). After fusing the features of each band through feature mapping, they are sent to the decoder.

[0096] The attention mechanism of this network is realized by introducing channel attention and spatial attention modules. The calculation formula for channel attention is:

[0097] M c = σ(MLP(AvgPool(F img )) + MLP(MaxPool(F img )));

[0098] where M c represents the channel attention weight matrix, which is used to weight the importance of different channel features. F img is the input feature, AvgPool and MaxPool are the average pooling and maximum pooling operations respectively, MLP is a multi-layer perceptron, and σ is the Sigmoid activation function.

[0099] In practical application scenarios, the network performs excellently when processing underwater images in the South China Sea with a turbidity of up to 30 NTU (turbidity unit). For example, for a set of PET plastic bag images taken at a depth of 5 meters underwater with severe blue-green shift and low contrast, after being processed by the network, the detail clarity (evaluated by Laplacian gradient) is improved by 285%, and the color restoration degree (evaluated by UCIQE index) is improved by 176%. This enables the subsequent detection algorithm to correctly identify partially visible plastic bags with an occlusion rate of 65%, while the detection algorithm on the original image cannot locate the target at all.

[0100] The image enhancement process is expressed as:

[0101] I enhanced = G gen (I raw , λ spec ) ;

[0102] Among them, I raw is the original underwater multi-spectral image, λ spec is the spectral wavelength parameter, G gen is the generation network function, and I enanced is the enhanced image.

[0103] The network optimization objective is to minimize the weighted sum of the generation loss and the adversarial loss:

[0104] L img_total = α img L recon + β img L adv + γ img L perceptual ;

[0105] Among them, L img_total represents the total loss function of the image enhancement network, L recon is the reconstruction loss, L adv is the adversarial loss, L perceptual is the perceptual loss, and α img , β img and γ img are weight coefficients.

[0106] The specific implementation of the multi-spectral image enhancement network adopts a generative adversarial network architecture:

[0107] G network (I raw ) = I enhanced ;

[0108] D network (I) ∈ [0,1];

[0109] Among them, Gnetwork is a generator network for converting the original underwater image I raw into the enhanced image I enhanced ; D network is a discriminator network that outputs an image authenticity score, and I is the input image.

[0110] Sub-step 2.2: Construct a 3D voxelization representation system based on binocular stereo vision;

[0111] Construct a voxelization representation system that uses binocular cameras for depth estimation and three-dimensional reconstruction, processes the binocular image sequence, and outputs the three-dimensional voxel representation of the objects in the scene. This system includes: a stereo matching module: calculates the disparity map and estimates the depth information; a point cloud generation module: converts the depth information into a three-dimensional point cloud; a voxelization module: voxelizes the point cloud data into a regular three-dimensional grid representation.

[0112] The voxelization representation process can be expressed as:

[0113] V = Ψ(P(D depth (I left , I right ))) ;

[0114] V = Ψ(P(D(I left , I right ))) ;

[0115] where I left and I right are the left and right camera images respectively, D depth represents the depth estimation function for calculating the depth information of the binocular images, D is the depth estimation function, P is the point cloud generation function, Ψ is the voxelization function, and V is the final voxel representation. This representation is particularly suitable for handling partial occlusion situations because it preserves the complete three-dimensional spatial information of the object.

[0116] The 3D voxelization representation system obtains the depth information through binocular stereo vision and constructs the voxel grid V:

[0117]

[0118] where (x, y, z) represents a point in three-dimensional space.

[0119] The voxelization process includes: binocular image matching to obtain the disparity map; converting the disparity map to the depth map; reconstructing the depth map into a point cloud; and voxelizing the point cloud data into a regular three-dimensional grid.

[0120] Sub-step 2.3: Construct a multi-scale feature fusion network;

[0121] Construct a multi-scale feature fusion network based on a feature pyramid to process image features at different scales and output a fused discriminative feature map. This network adopts an adaptive weight mechanism to dynamically adjust the fusion weights according to the importance of features at different scales.

[0122] The feature fusion process is expressed as:

[0123]

[0124] Among them, represents the summation operation from i = 1 to n, where i is the summation variable and n is the upper limit of the summation. F i is the feature map of the i-th scale, W i is the corresponding transformation matrix, α i is the adaptive weight coefficient, and F fused is the fused feature map.

[0125] The weight coefficient is generated through an attention mechanism:

[0126]

[0127] Among them, F i represents the feature map of the i-th scale, F j represents the feature map of the j-th scale, exp represents the natural exponential function, represents the summation operation from j = 1 to n, where j is the summation variable and n is the upper limit of the summation. A attn is the attention evaluation function, which evaluates the importance of the feature map for the current object detection.

[0128] Sub-step 2.4: Construct a motion plastic tracking system based on temporal attention;

[0129] Construct an attention mechanism that integrates temporal information to process a sequence of consecutive frame images and output stable plastic object tracking results. This system enhances the tracking ability of moving plastics in water by establishing spatio-temporal correlations between consecutive frames.

[0130] The temporal attention calculation is expressed as:

[0131] A t = softmax(f query (F t )·f key (F t-k:t-1 ) T )·f value (F t-k:t-1 ));

[0132] Among them, softmax represents the softmax function, which is used to convert a vector into a probability distribution. A tis the attention feature at time step t, F t is the feature at the current time step, F t-k:t-1 are the features of the previous k time steps, f query 、f key and f value are the query, key, and value mapping functions respectively, and T represents the matrix transpose operator.

[0133] Step 3: Collect plastic feature data through multispectral imaging technology, and combine with a multi-task learning framework to simultaneously perform PET material type identification, degradation degree assessment, and morphological analysis;

[0134] In this step, by constructing a feature extraction and classification system for the fouling characteristics of the marine environment, the problem of difficult identification caused by biofilms, salts, and pollutants attached to the surface of marine PET plastics is solved. It specifically includes the following sub-steps:

[0135] Sub-step 3.1: Construct a multispectral imaging data acquisition and preprocessing system;

[0136] Construct a multispectral imaging system covering the 300 - 2500 nm band, collect multi-band image data of marine PET plastics, and perform standardization processing to output a corrected multispectral data cube. This system includes: a multispectral sensor array: collect image data of different bands; a spectral correction module: correct the light intensity difference between different bands; a geometric registration module: ensure the spatial alignment of images of different bands.

[0137] The multispectral data cube is expressed as:

[0138] I(x,y,λ) = C corr (R(x,y,λ));

[0139] where R(x,y,λ) is the original spectral response at wavelength λ at position (x,y), C corr is the correction function, and I(x,y,λ) is the corrected multispectral data cube.

[0140] Sub-step 3.2: Construct a dual-threshold adaptive segmentation algorithm;

[0141] Construct a dual-threshold adaptive segmentation algorithm for complex backgrounds, process the corrected multispectral images, and output the segmentation result of the PET plastic region after removing interference factors. This algorithm automatically selects the most discriminative band combination through multispectral characteristic analysis to effectively eliminate interference factors such as biofilms and salts.

[0142] The adaptive segmentation process is expressed as:

[0143]

[0144] Among them, I(x, y, λ1), I(x, y, λ2), I(x, y, λ k ) represent the spectral response values at position (x, y) with wavelengths λ1, λ2, λ k , respectively. These values together constitute a sequence of response values in different bands at a certain spatial position in the hyperspectral data cube. F band is a band combination function, and λ1, λ2,..., λ k are the selected most discriminative bands. T low and T high are the adaptively determined lower and upper thresholds, and M(x, y) is a binary mask indicating whether the position (x, y) is a PET plastic area. The thresholds are automatically updated by combining the Otsu algorithm with historical data statistics:

[0145] T low , T high = O(H(F band (I)));

[0146] Among them, H is a histogram function, O is an Otsu threshold calculation function based on maximizing the between-class variance, and I represents the corrected hyperspectral data cube.

[0147] Sub-step 3.3: Construct a multi-task learning framework;

[0148] Construct a multi-task learning framework that integrates PET material type recognition, degradation degree assessment, and morphological analysis, processes the segmented PET plastic area images, and outputs the prediction results of three key attributes simultaneously. This framework adopts a shared encoder - multi-task decoder architecture, making full use of knowledge sharing and transfer between different tasks.

[0149] This multi-task learning framework adopts a backbone network structure with hard parameter sharing, uses ResNet-50 as the shared encoder for feature extraction, and configures dedicated decoders for three different tasks respectively. The shared encoder contains 5 residual blocks, and each residual block contains multiple convolutional layers, batch normalization layers, and skip connections. The three task-specific decoders each contain two fully connected layers. Among them, the last layer of the material type recognition decoder uses the Softmax activation function (outputting the probability distribution of 6 common marine PET types), the degradation degree assessment decoder uses the Sigmoid activation function (outputting a continuous value between 0 and 1, representing the degradation degree percentage), and the morphological analysis decoder uses the linear activation function (outputting continuous values representing physical parameters such as size and thickness).

[0150] The framework introduces a task balance mechanism to avoid the problem of task dominance during the training process by dynamically adjusting the loss weights of each task:

[0151]

[0152] Among them, represents the loss value of the i-th task in the t-th batch, and λ temp is the temperature parameter that controls the smoothness of weight allocation. exp() represents the natural exponential function, represents the weight of the i-th task in the t-th batch, represents the summation over 3 tasks.

[0153] The overall loss function of the multi-task learning framework is:

[0154]

[0155] Among them, and represent the loss functions of material type recognition, degradation degree evaluation, and morphological analysis respectively. λ1, λ2, and λ3 are the weight coefficients of each task. The weight coefficients are dynamically adjusted according to the performance on the validation set.

[0156] In practical applications, this framework can handle multiple attribute prediction requirements simultaneously. For example, for a batch of PET beverage bottles recovered from the offshore area of Hainan Island, the system can simultaneously complete the following analyses: 1) accurately identify that it is a polyethylene terephthalate beverage bottle (confidence level 98.3%); 2) evaluate its degradation degree as 37.2%, which is in the medium degradation stage; 3) analyze that its structural integrity is 83.5% and the thickness wear deviation is 0.12 mm. Based on this multi-dimensional information, the recycling processing system automatically assigns it to the processing process suitable for making recycled fibers, rather than the plastic pellet reprocessing process that requires higher material integrity, improving the resource utilization efficiency. In experimental verification, compared with the method of training single-task models separately and then integrating them, the inference speed of this multi-task framework is increased by 2.8 times, the number of parameters is reduced by 65%, and the average prediction accuracy is increased by 7.6%.

[0157] The multi-task learning process is expressed as:

[0158] y type , y deg , y morph = Dec type (E encoder (x)), Dec deg (E encoder (x)), Dec morph (E encoder (x));

[0159] Among them, x is the input image, E encoder is the shared feature encoder, Dec type , Dec deg and Dec morphSpecial decoders for material type, degradation degree, and morphological analysis, respectively, y type 、y deg and y morph are the output results of the three tasks respectively. The optimization objective is the weighted task loss:

[0160] L MT_total = w type L type + w deg L deg + w morph L morph ;

[0161] where L MT_total represents the total loss function of multi-task learning, which is the weighted sum of the loss functions of the three tasks. L type 、L deg and L morph are the loss functions of the three tasks respectively, and w type 、w deg and w morph are the corresponding weight coefficients.

[0162] Sub-step 3.4: Construct a deep learning model for marine environmental fouling;

[0163] Construct a deep learning model that adapts to the characteristics of marine environmental fouling, processes the preprocessed PET plastic images, and outputs the final recognition and classification results. This model uses the transfer learning method to fine-tune using a pre-trained network and a specific marine fouling dataset to improve its adaptability to complex fouling conditions.

[0164] The model construction process includes the following links: Selection of pre-trained network: Select a backbone network pre-trained on a large-scale dataset; Configuration of domain adaptation layer: Add a dedicated domain adaptation layer to reduce the difference in feature distributions between the source domain and the target domain; Implementation of multi-instance learning: Make a joint decision through multiple perspectives of the same sample to improve robustness.

[0165] The domain adaptation process is expressed as:

[0166]

[0167] where M s and M t are the feature mapping functions of the source domain and the target domain respectively, x s and x t are the samples of the source domain and the target domain respectively, ||·|| F is the Frobenius norm, and D loss is the domain adaptation loss, and the domain difference is reduced by minimizing this loss.

[0168] Step 4: The multi-agent collaborative grasping control system based on reinforcement learning realizes the precise recycling of PET plastics;

[0169] In this step, by constructing a multi-agent collaborative control system based on reinforcement learning, the multi-agent system of sea, land and air, the underwater PET plastic multi-object detection model, and the fouled PET plastic feature extraction and classification system are integrated. This system optimizes the collaborative grasping strategy of multi-agents by maximizing the cumulative reward function Q(s, a), where the state space s contains information such as the positions of each agent and the positions of target PET plastics, and the action space a contains the movement and grasping actions of the agents. The system uses the Deep Q-Network (DQN) for policy learning, and improves the learning stability through techniques such as experience replay and target network, and finally forms a complete marine recycled PET plastic recycling and processing system. Specifically, it includes the following sub-steps:

[0170] Sub-step 4.1: Construct a distributed microservice architecture;

[0171] Construct a distributed microservice architecture based on container technology to manage the communication and data interaction between subsystems, and achieve high availability and scalability of the system. This architecture includes: Service Registration and Discovery Center: Manage the addresses and status of each microservice; API Gateway: Unified interface management and request routing; Message Queue: Achieve asynchronous communication between systems; Data Storage Service: Manage the storage and retrieval of various types of data.

[0172] The communication between services adopts the event-driven mode, and realizes the loose-coupled system interaction through the message queue:

[0173] E events ={e1, e2,..., e n};

[0174] where, E events is the event set, and each service interacts by publishing and subscribing to events. e1, e2, e n represent the 1st, 2nd, and nth events respectively.

[0175] Sub-step 4.2: Construct a multi-level decision fusion mechanism;

[0176] Construct a multi-level decision fusion mechanism that combines a rule engine and a learning-based decision model to comprehensively analyze the decision suggestions from different subsystems and output the final coordinated system behavior instructions. This mechanism includes: Rule Engine: Encode expert knowledge and safety constraints; Bayesian Decision Network: Process decision optimization under uncertain conditions; Conflict Resolution Module: Coordinate decision conflicts between different subsystems.

[0177] The multi-level decision-making fusion mechanism adopts a four-layer architecture: the first layer is the rule-based reasoning layer, implemented by the Drools rule engine, which contains about 200 expert knowledge rules covering fixed logics such as marine environmental safety constraints, equipment operation restrictions, and emergency handling; the second layer is the deep reinforcement learning layer, which adopts an improved DQN network structure to learn the optimal decision-making strategy through interaction with the environment; the third layer is the probabilistic reasoning layer, which adopts a dynamic Bayesian network structure, containing 35 nodes and 52 edges, representing the temporal probability relationships among decision-making factors; the fourth layer is the multi-objective optimization layer, which achieves the balance of multiple objectives such as efficiency, safety, resource consumption, and environmental impact through the genetic algorithm.

[0178] The multi-agent collaborative grasping control system based on reinforcement learning optimizes the strategy by maximizing the cumulative reward function:

[0179]

[0180] Among them, π represents the policy function, represents the cumulative summation from time step 0 to T RL , E represents the expected value operator, argmax represents the variable value corresponding to the maximum value of the objective function, π * is the optimal policy, r t is the reward at time step t, represents the discount factor at time t, T RL is the task completion time. This optimization process ensures the collaborative efficiency of the multi-agent system in a complex marine environment.

[0181] The conditional probability table of the Bayesian network is obtained by learning historical data and updating parameters using the maximum likelihood estimation and expectation maximization algorithms:

[0182]

[0183] Among them, represents the conditional probability estimation that node i takes value k under the state j of the parent node, N ijk is the corresponding observation count, α ijk is the prior pseudo count, ∑ k represents the summation over all possible values k.

[0184] In practical applications, this mechanism can effectively handle multi-system conflict decision-making situations. For example, in a large-scale ocean plastic floating belt recovery mission, the identification system detected a large accumulation of PET plastics and recommended dispatching all unmanned boats to quickly go for recovery; at the same time, the safety monitoring system detected an upcoming thunderstorm and recommended that all equipment retreat to avoid danger; while the resource optimization system recommended selectively dispatching some unmanned boats to perform tasks according to the remaining battery power. Facing such conflicting suggestions from multiple parties, the decision fusion mechanism first determines through the rule engine that thunderstorm safety is the highest priority constraint, then uses the Bayesian network to evaluate the risk probability (determining that the thunderstorm will actually arrive in 2 hours), and finally generates a compromise solution through a multi-objective optimization algorithm: immediately dispatch two unmanned boats with the highest remaining battery power to perform rapid recovery, set a mandatory return limit of 90 minutes, and at the same time other equipment stays in place as backup. This decision not only meets the safety constraints but also achieves partial recovery goals, which is better than the decisions of any single system.

[0185] The decision fusion process is expressed as:

[0186]

[0187] where argmax represents the value of the variable corresponding to when the objective function reaches the maximum value, d * represents the optimal fusion decision, that is, the finally selected decision plan, d is the alternative decision, D space is the decision space, U i is the utility function of the i-th subsystem, C env is the current environmental condition, is the subsystem weight, d * is the optimal fusion decision, represents the sum of all terms from i = 1 to m, where m is the total number of subsystems.

[0188] Sub-step 4.3: Construct a human-machine collaborative interaction interface;

[0189] Construct a human-machine collaborative interaction interface that integrates visual display and instruction input, supports human operators to monitor the system status and perform necessary interventions, and ensures the safe and effective operation of the system in complex environments. This interface includes: multi-view display: showing the key status and detection results of each subsystem; 3D scene reconstruction: visualizing the current sea area environment and target distribution; hierarchical control mode: supporting the switching of multi-level control modes from fully automatic to fully manual.

[0190] The interaction interface adopts an adaptive information display strategy and dynamically adjusts the display content according to the current situation and the operator's focus:

[0191] I d = f(S, C UI , A user );

[0192] Among them, S is the system state, C UI is the current environmental condition, A user is the operator's attention focus, I d is the optimized information display content, and f represents the information display optimization function, which is used to dynamically adjust the display content according to the system state, environmental conditions, and operator attention.

[0193] Sub-step 4.4: Build a system self-diagnosis and optimization module;

[0194] Build a self-diagnosis and optimization module based on operation data analysis, continuously monitor the system operation state, automatically detect potential problems and adjust parameters to maintain the optimal performance of the system. This module includes: Anomaly detector: Identify abnormal behaviors of system components; Performance analyzer: Evaluate the operation efficiency of each subsystem; Parameter optimizer: Automatically adjust system parameters according to operation data.

[0195] The system optimization process is expressed as:

[0196]

[0197] Among them, θ sys is the system parameter set, E env is the environmental condition, S(θ sys , E env ) is the system performance under given parameters and environment, L perf is the performance loss function, is the optimized parameter set, and argmin represents the variable value corresponding to the minimum value of the objective function.

[0198] Through the innovative integration of a multi-agent system with sea, land, and air collaboration, a multi-objective detection model adapted to the underwater environment, and a feature extraction and classification system for fouling characteristics, this embodiment realizes the all-terrain active source search, precise identification, and automatic recycling and treatment of marine PET plastics. This method has the following remarkable technical effects:

[0199] It breaks through the limitation of single-environment recognition, realizes the ability to recycle and process PET plastics across media and all terrains, and the recycling efficiency is increased by 3.5 times, which is especially suitable for the efficient recycling of PET plastics in complex environments such as coastlines and near islands.

[0200] It solves the influence of the complex underwater optical environment on recognition. The system can directly identify underwater PET plastics without changing the marine environment, realizes the real-time and in-situ recognition of underwater PET plastics, and greatly improves the efficiency and coverage of marine PET plastic recycling.

[0201] Overcoming the interference of marine fouling on PET plastic identification, the system achieves an identification accuracy of 94.7% on PET samples collected in the real marine environment, which is 28.3% higher than that of traditional RGB image recognition systems, and can operate stably in high-salt and highly fouled environments.

[0202] Through advanced multi-scale feature fusion and temporal attention mechanism, this method significantly improves the identification ability of PET plastics in the case of partial occlusion and multi-object stacking, with the accuracy rate increased by 31.5%, providing a reliable basis for the fine classification of various types of plastics.

[0203] The innovative multi-task learning framework enables the system to simultaneously evaluate the PET material type, degradation degree and morphological characteristics, providing more comprehensive material information support for subsequent processing and reuse, and improving the utilization value of recycled PET plastics.

[0204] Real application examples of this embodiment

[0205] 1. Application scenario description: This embodiment is applied to a PET plastic pollution control project in a certain sea area of the South China Sea. Due to the developed surrounding tourism and frequent marine fishery activities in this sea area, there are a large number of PET plastic wastes. The treatment area covers three typical environments: open water area (water depth 5 - 25 meters), multi-island reef environment area (water depth 3 - 15 meters, with coral reefs and rocky areas), and nearshore beach area (water depth 0 - 5 meters). The seawater turbidity varies from 5 to 40 NTU, there are phytoplankton and debris on the water surface, and the underwater light attenuation is obvious. The main sources of PET plastics are beverage bottles, packaging bags, fishing gear, etc., which are distributed on the water surface, underwater and on the shore. Some of them have biological attachment and salt erosion, and there are obvious deteriorations in appearance.

[0206] The technical challenges in the application scenario include: difficulties in identification caused by underwater light scattering, confusion in the identification of multiple types of plastics, partial occlusion of some plastics by seagrass or corals, insufficient evaluation of the degradation degree, etc., making the existing recycling systems inefficient and with a high loss rate.

[0207] 2. Application examples of the embodiment

[0208] 2.1. Deployment of multi-agent system:

[0209] A multi-agent system composed of the following agents is deployed in the South China Sea test sea area: 3 surface unmanned boats: equipped with multi-spectral cameras, GPS navigation systems, and grasping robotic arms; 2 underwater robots: equipped with underwater multi-spectral imaging systems, sonar detectors, and robotic grippers; 1 aerial drone: equipped with high-resolution cameras and infrared imagers.

[0210] The following table shows the configuration parameters of the multi-agent system:

[0211]

[0212] The multi-agent collaborative behavior tree includes the following main nodes: global task allocation (selecting sub-regions, allocating search tasks); environmental perception and fusion (summarizing multi-source data, constructing an environmental model); target detection and tracking (identifying the location of plastics, dynamically tracking); collaborative capture strategy (multi-agent path planning, grasping division of labor); recycling and transportation (classified storage, returning and unloading).

[0213] 2.2 Application of self-organizing map prediction model:

[0214] The system collected historical data in this sea area over the past two years, including wind direction, tide, ocean current data, and records of the locations where plastics were found in the past. Based on these data, a self-organizing map of 7×7×3 (spatial resolution) × 24 (time resolution) was constructed, and the sea area was divided into 504 spatio-temporal units for density prediction.

[0215] The following table shows the accuracy evaluation of the self-organizing map prediction model:

[0216]

[0217] One actual application case: After the typhoon "Goni", the sea area was detected. The system analyzed the wind direction (southeasterly, level 6), typhoon path, tide state, and historical data, and predicted 5 high-probability plastic aggregation areas. The search results showed that a large amount of PET plastics aggregated in 4 of these areas. Especially in the sheltered bay on the northeast side of the island reef, more than 300 kg of PET plastic waste was found, and the actual capture weight was 4.8 times that of the conventional grid search method.

[0218] 2.3 Application of underwater multi-spectral image enhancement and recognition:

[0219] In an underwater turbid environment (turbidity 28 NTU), the underwater PET plastic images are processed using the multi-spectral image enhancement network of this embodiment. The system collected 10-band images in the range of 400 - 900 nm, and after enhancement processing, a comparison was made with traditional image enhancement methods.

[0220] The following table shows the comparison of the multi-spectral image enhancement effects:

[0221]

[0222] Application Case: At a depth of 8 meters underwater in a reef area, a traditional RGB camera was completely unable to identify a transparent PET beverage bottle that was partially blocked by coral (blocking rate approximately 45%). Using the multi-spectral image enhancement system of this embodiment, after acquiring images at two characteristic bands of 850 nm and 490 nm and performing enhancement, the system successfully separated the PET plastic from the complex background and accurately identified it as a transparent PET beverage bottle. Even with a 45% blocking rate, the identification accuracy still reached 92.3%.

[0223] After image enhancement, the system uses 3D voxelization representation technology to reconstruct the complete shape of the occluded object. In 50 random occlusion tests, the average reconstruction accuracy reached 87.5%, far higher than 63.2% of the traditional 2D method.

[0224] 2.4. Application of Multi-Task Learning Framework:

[0225] The system collected 600 marine PET plastic samples under different conditions, covering 6 common types, 5 degradation levels, and various morphological characteristics. The samples were divided into a training set, a validation set, and a test set in a ratio of 7:1:2 to train a multi-task learning model and compare it with a separately trained model.

[0226] The following table shows the performance comparison between the multi-task learning framework and the single-task model:

[0227] Performance indicators Single - task model (trained separately) Multi - task learning model Performance improvement Number of parameters (in millions) 118.5 41.7 Reduced by 64.8% Inference speed (samples / second) 7.8 22.1 Improved by 183.3% Accuracy of material type recognition 89.3% 94.7% Improved by 5.4% Error in degradation degree evaluation 12.8% 8.5% Reduced by 33.6% Accuracy of morphological analysis 81.5% 90.2% Improved by 8.7% Average performance Benchmark - Improved by 15.9%

[0228] Application Case: In a recycling operation, the system recovered a batch of PET plastic fragments (about 25 pieces, with an area between 10 - 500 cm 2 from a depth of 5 meters underwater). The multi-task learning framework analyzed each piece simultaneously, and the results showed: Material identification: 23 pieces were beverage bottle-grade PET materials (confidence > 95%), and 2 pieces were food packaging-grade PET materials (confidence > 92%); Degradation assessment: The degradation level ranged from 22% - 65%, with an average of 42.7%; Morphological analysis: The average thickness was about 0.31 mm, and the surface roughness increased by 187% compared to new materials.

[0229] Based on this multi-dimensional information, the system automatically assigned the fragments with a degradation degree < 40% to the recycled plastic pellet process and the fragments with a degradation degree > 40% to the fiber manufacturing process, optimizing the material reuse value. Physical detection verification showed that the consistency rate between the processing plan assigned by the system and the manual expert evaluation reached 94.3%.

[0230] 2.5. Application of Multi-Agent Cooperative Grasping:

[0231] Under actual sea conditions (medium waves, water flow speed 0.8 m / s), the system targeted large PET plastic garbage (area about 3.2 m 2The collaborative grasping test was carried out to compare the effects of single-agent and multi-agent collaborative schemes.

[0232] The following table shows the evaluation of the collaborative grasping effect of multi-agents:

[0233] Performance indicators Single - task model (trained separately) Multi - task learning model Performance improvement Number of parameters (in millions) 118.5 41.7 Reduced by 64.8% Inference speed (samples / second) 7.8 22.1 Improved by 183.3% Accuracy of material type recognition 89.3% 94.7% Improved by 5.4% Error in degradation degree evaluation 12.8% 8.5% Reduced by 33.6% Accuracy of morphological analysis 81.5% 90.2% Improved by 8.7% Average performance Benchmark - Improved by 15.9%

[0234] Application case: Recycling a PET plastic fishing net with an area of about 2.8 m 2 in a reef area in the South China Sea. The system automatically coordinated the following collaborative behaviors: 1. The aerial drone first conducted high-altitude surveys to determine the target position and shape; 2. Two surface unmanned boats approached from the upstream and the side respectively to form a "V"-shaped encirclement formation; 3. The underwater robot approached from below to provide an underwater perspective and support; 4. The collaborative control system based on reinforcement learning real-time allocated the grasping positions and sequences of each agent; 5. First, the underwater robot fixed the bottom of the fishing net to prevent it from sinking, then the two surface unmanned boats successively grasped the left and right ends of the net, and finally cooperated to curl and package it.

[0235] The whole process was completed within 98 seconds with a success rate of 100%, while 5 attempts by a single surface unmanned boat under the same conditions all failed.

[0236] 3. Verification of technical effects

[0237] Compared with the prior art, this embodiment has achieved significant improvements in two key technical effects: recognition accuracy and recovery efficiency. The following is verified by the actual test data in the South China Sea experimental area.

[0238] The following table shows the comparison of plastic recognition accuracy under different environmental conditions:

[0239]

[0240] The following table shows the comparison of PET plastic recovery efficiency under different environments:

[0241] Environment type Recovery method Recovery amount (kg / hour) Unit energy consumption (Wh / kg) Cost - benefit ratio Open water Manual recovery 5.3 - 1.0 Open water Traditional vessel 12.7 423 2.4 Open water This embodiment 42.6 86 8.0 Reef area Manual recovery 3.1 - 1.0 Reef area Traditional vessel 4.5 612 1.5 Reef area This embodiment 23.8 124 7.7 Near - shore shallow water Manual recovery 7.6 - 1.0 Near - shore shallow water Traditional vessel 8.2 378 1.1

[0242] Through six months of actual application tests, this embodiment has recycled 8.7 tons of PET plastic in the South China Sea test area. Compared with the expected recovery amount of 2.4 tons by the traditional method under the same conditions, the improvement rate has reached 262.5%. The system especially performs outstandingly in complex environments, and the recovery efficiency in the reef area has increased by 4.3 times, solving the technical problems that are difficult for traditional systems to handle.

[0243] The multi-agent collaborative control system can still maintain an 83.6% recognition accuracy and a 76.2% recovery success rate under extreme weather conditions (wind force level 6, wave height 1.5 m), far exceeding the performance of traditional systems under the same conditions (32.1% and 21.5% respectively).

[0244] Experimental data shows that the two core technical indicators of this embodiment have achieved the expected goals: recognition accuracy: on average, it reaches 92.2% in complex marine environments, an increase of 71.7% compared to traditional technologies; recovery efficiency: the comprehensive efficiency reaches 31.6 kg / hour, which is 3.7 times that of traditional methods, and the energy consumption is reduced by 78%.

[0245] In addition, the system evaluates the degradation degree and material properties of recycled PET plastics through a multi-task learning framework, enabling the classification accuracy of subsequent recycling treatment to reach 94.3%, greatly improving the resource utilization value of marine recycled PET plastics.

[0246] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A marine recycled PET plastic recycling and treatment process, characterized in that, Including the following steps: Construct a multi-agent system through surface unmanned vessels, underwater robots and aerial unmanned aircraft, conduct collaborative decision-making and task allocation based on the behavior tree framework, and use the self-organizing map model to predict the distribution density of PET plastics to guide the multi-agent system for efficient source seeking; Utilize the multi-spectral image enhancement network and self-attention detection architecture for the underwater environment, combine binocular stereo vision to construct a 3D voxel representation, and achieve accurate identification of underwater PET plastics; Collect plastic feature data through multi-spectral imaging technology, and simultaneously conduct PET material type identification, degradation degree evaluation and morphological analysis in combination with the multi-task learning framework; The multi-agent collaborative grasping control system based on reinforcement learning realizes the accurate recovery of PET plastics.

2. The marine recycled PET plastic recycling and treatment process according to claim 1, characterized in that, The multi-agent system adopts a cross-media perception information fusion algorithm to fuse and process multi-source data from the air, surface and underwater, and the fusion process is expressed as: D fused = f fusion (D air , D surface , D underwater ); Among them, D fused represents the fused perception data, D air , D surface and D underwater respectively represent the aerial, surface, and underwater perception data, and f fusion is the fusion function.

3. The marine recycled PET plastic recycling process according to claim 1, characterized in that, The self-organizing map model calculates the distribution density of PET plastics through the following steps: Among them, ρ(x, y, z, t) is the predicted value of the PET plastic density at the spatial point (x, y, z) at time t, represents the weighted sum of k basis functions, φ i (x, y, z, t) represents the i-th basis function, which is used to describe the space of plastic distribution, is the corresponding weight coefficient, k is the total number of basis functions, and the weight coefficient is continuously updated through an online learning method to adapt to the changing marine environment.

4. The marine recycled PET plastic recycling and treatment process according to claim 1, characterized in that, The multi-spectral image enhancement network adopts a generative adversarial network architecture, including: G network (I raw ) = I enhanced ; D network (I) ∈ [0, 1]; Among them, G network is the generator network, I raw is the original underwater image, I enhanced is the enhanced image, D network is the discriminator network, which outputs the authenticity score of the image, and I is the input image.

5. A marine recycled PET plastic recycling and treatment process according to claim 1, characterized in that, The 3D voxel representation obtains depth information through binocular stereo vision and constructs a voxel grid V: Where (x, y, z) represents a point in three-dimensional space.

6. The marine recycled PET plastic recycling and treatment process according to claim 1, wherein, The multi-task learning framework simultaneously optimizes the loss functions of three tasks: Among them, and represent the loss functions of material type recognition, degradation degree evaluation, and morphological analysis respectively, and λ1, λ2, and λ3 are the weight coefficients of each task.

7. A marine recycled PET plastic recycling and treatment process according to claim 1, characterized in that, The multi-agent collaborative grasping control system based on reinforcement learning optimizes the policy by maximizing the cumulative reward function: Among them, π represents the policy function, represents the cumulative sum from time step 0 to T RL , E represents the expected value operator, argmax represents the variable value corresponding to the maximum value of the objective function, and π * is the optimal policy, r t is the reward at time step t, represents the discount factor at time t, and T RL is the task completion time, and this optimization process ensures the cooperation efficiency of the multi-agent system in the complex marine environment.

8. A marine recycled PET plastic recycling process according to claim 1, characterized in that, The multi-scale feature fusion mechanism is calculated according to the following formula: Among them, represents the summation operation from i = 1 to n, where i is the summation variable and n is the upper limit of the summation, and F i is the feature map of the i-th scale, and W i is the corresponding transformation matrix, and α i is the adaptive weight coefficient, and F fused is the fused feature map.

9. A marine recycled PET plastic recycling and treatment process according to claim 1, characterized in that, It also includes a temporal attention mechanism step for capturing the motion characteristics of plastics in the water flow: A t = softmax(f query (F t )·f key (F t-k:t-1 ) T )·f value (F t-k:t-1 ); Among them, softmax represents the softmax function, which is used to convert a vector into a probability distribution, and A t is the attention feature at time step t, and F t is the feature at the current time step, and F t-k:t-1 is the features of the previous k time steps, and f query , f key and f value are the query, key, and value mapping functions respectively, and T represents the matrix transpose operator.

10. A marine recycled PET plastic recycling and treatment system for implementing the marine recycled PET plastic recycling and treatment process described in any one of claims 1-9, characterized in that, Including: A multi-spectral imaging module for collecting multi-band spectral feature data of underwater PET plastics; A GPS positioning module for obtaining the accurate position information of each agent in real time; A surface maneuver execution module responsible for the motion control of the surface unmanned vessel; A binocular stereo vision module for constructing a 3D voxel representation of underwater targets; An underwater propulsion and grasping module responsible for the motion and grasping actions of the underwater robot respectively; A flight control module responsible for the flight path planning and attitude control of the aerial unmanned aircraft; An image enhancement module for improving the quality of underwater images through a generative adversarial network; A target detection module for accurately positioning PET plastics based on the self-attention mechanism; A material analysis module for evaluating plastic properties using the multi-task learning framework; A multi-agent collaborative decision-making module for task allocation based on the behavior tree framework; A task planning module responsible for formulating global search and recovery strategies; A data processing module for fusing and analyzing multi-source sensing data.

Citation Information

Cited By

  • Water surface target grabbing system and control method thereof

    CN120863811A

  • A surface target acquisition system and its control method

    CN120863811B