Control method of robotic arm, sorting equipment and storage medium

By aligning and denoising multimodal data at the edge gateway, generating cargo features and generating sorting strategies, the inefficiency of integration caused by the difference in multimodal data formats is solved, and the sorting efficiency is improved.

CN120245020BActive Publication Date: 2025-08-12GUANGZHOU PINGYUN CRAFTSMAN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510758408.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-12
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

Large differences in multimodal data formats lead to low information integration efficiency, which leads to low sorting efficiency.

Method used

By performing time alignment and modal type noise reduction processing at the edge gateway, multimodal structured data is generated, and dynamic confidence is determined based on ambient light data and fused cargo characteristics are generated, combining cargo priority and equipment operation data to generate sorting strategies, and adjust the end position and jaw torque parameters of the robot arm.

Benefits of technology

It realizes efficient integration of multimodal data and improves sorting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120245020B_ABST
    Figure CN120245020B_ABST
Patent Text Reader

Abstract

This application discloses a control method for a robotic arm, a sorting device, and a storage medium. This application relates to the technical field of warehouse management systems. The control method for the robotic arm includes: performing time alignment on an edge gateway for sensor data collected by an image acquisition module, a radio frequency identification tag, and an environmental sensor to obtain time-synchronized data; performing noise reduction processing on the time-synchronized data based on modality type to obtain multimodal structured data; fusing the multimodal structured data to obtain cargo features based on dynamic confidence determined by ambient light data; inputting the cargo features into a sorting engine, combining cargo priority and equipment operation data to generate a sorting strategy; and adjusting the end-of-line posture of the robotic arm and the torque parameters of the gripper based on the sorting strategy. This method achieves the technical effect of integrating multimodal data to improve sorting efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of warehouse management systems, and in particular to a control method for a robotic arm, a sorting device, and a storage medium. Background Art

[0002] In the after-sales service field of smart devices, sorting devices are often installed to enable the sequential sorting and delivery of goods. Related technologies further deploy AI (artificial intelligence) cameras, RFID (radio frequency identification) tags, temperature and humidity sensors, and weight sensors to collect real-time cargo information and environmental data. This leverages AI visual recognition, the Internet of Things, and the integration of automated equipment to improve the accuracy and efficiency of warehouse operations.

[0003] However, the multimodal data formats in related technologies vary greatly, resulting in low information integration efficiency and, in turn, low sorting efficiency. Summary of the Invention

[0004] The main purpose of this application is to provide a control method, sorting equipment and storage medium for a robotic arm, aiming to solve the technical problem that the large differences in multimodal data formats lead to low information integration efficiency and thus low sorting efficiency.

[0005] To achieve the above objectives, the present application provides a method for controlling a robotic arm, the method comprising:

[0006] The sensor data collected by the image acquisition module, RFID tags, and environmental sensors are time-aligned at the edge gateway to obtain time-synchronized data;

[0007] Performing noise reduction processing on the time synchronization data based on the modality type to obtain multimodal structured data;

[0008] fusing the multimodal structured data to obtain cargo features based on a dynamic confidence level determined by the ambient light data;

[0009] Input the cargo characteristics into the sorting engine, combine the cargo priority and equipment operation data, and generate a sorting strategy;

[0010] The end speed of the robot arm and the torque parameters of the gripper are adjusted based on the sorting strategy.

[0011] In one embodiment, the step of time-aligning the sensor data collected by the image acquisition module, the radio frequency identification tag, and the environmental sensor at the edge gateway to obtain time-synchronized data includes:

[0012] receiving the image stream captured by the image acquisition module, the RFID data captured by the RFID tag, and the environmental data captured by the environmental sensor;

[0013] interpolating the image stream based on a cubic spline according to a frame rate of the image stream and an acquisition frequency of the RFID data to align the image stream and the RFID data;

[0014] Based on the environmental data at each time stamp, and the image stream and the RFID data aligned at the time stamp, it is marked as the time synchronization data.

[0015] In one embodiment, the step of performing noise reduction processing on the time synchronization data based on the modality type to obtain multimodal structured data includes:

[0016] performing noise reduction processing on the time synchronization data based on the modality type of the time synchronization data;

[0017] Determine abnormal data in the time synchronization data based on the rule engine and the equipment operation data in the warehouse;

[0018] The denoising result is filtered to remove the abnormal data and then encapsulated into the multimodal structured data.

[0019] In one embodiment, the time synchronization data further includes gyroscope data and device operation data, the device operation data including the speed of the end of the robot arm and the opening and closing angle of the gripper. The step of performing noise reduction processing on the time synchronization data based on the modal type of the time synchronization data includes any one of the following:

[0020] Decomposing the gyroscope data by three-level wavelet decomposition to obtain a position signal with a preset accuracy;

[0021] The equipment operation data is subjected to dynamic compensation for random drift errors, the Labuda criterion is combined to eliminate singular values, and the trend term is removed by the least square method;

[0022] Pass the image stream through adaptive histogram equalization.

[0023] In one embodiment, the abnormal data includes format error data and out-of-limit data, and the step of determining the abnormal data in the time synchronization data based on the rule engine and the equipment operation data in the warehouse includes:

[0024] determining, based on the rules engine, formatted data in the time synchronization data;

[0025] Building an isolation tree based on the equipment operation data, cargo weight, and each of the time synchronization data;

[0026] Out-of-limit data is determined based on the path length and anomaly score of each data point in the isolation tree.

[0027] In one embodiment, before the step of fusing the multimodal structured data with the dynamic confidence determined based on the ambient light data to obtain cargo characteristics, the following steps are included:

[0028] The visual feature vector corresponding to the image stream is extracted through the feature extraction module; wherein, the data input format of the feature extraction module is industrial resolution; the first layer of the feature extraction module is a 5×5 large convolution kernel, and the subsequent layers use a 3×3 small convolution kernel; each layer is connected and fused with a 128-channel conv3_x, a 256-channel cnv4_x, and a 512-channel conv5_x, and then the fusion output is a 256-dimensional visual feature vector.

[0029] In one embodiment, the step of fusing the multimodal structured data with the dynamic confidence determined based on the ambient light data to obtain cargo characteristics includes:

[0030] Determining the dynamic confidence level by looking up a table based on the ambient light data;

[0031] The cargo feature is obtained by fusing the visual feature vector and the radio frequency identification data according to the dynamic confidence.

[0032] In one embodiment, the step of inputting the cargo characteristics into a sorting engine and generating a sorting strategy based on cargo priorities and equipment operation data includes:

[0033] splicing the cargo characteristics, the cargo priority parsed from the RFID data, and the device operation data into a state vector;

[0034] The state vector is input into the sorting engine, and the output of the sorting engine is used as the sorting strategy. The sorting engine includes a 256-dimensional input layer, a two-layer hidden layer with 128 units for processing time series data, and an output layer that discretizes 51 action branches, where the action branches include the end velocity of the robot arm and the opening and closing angle of the gripper.

[0035] In addition, to achieve the above-mentioned purpose, the present application also provides a sorting device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein the computer program is configured to implement the steps of the control method of the robotic arm as described above.

[0036] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer-readable storage medium, and the computer-readable storage medium stores a program for implementing the control method of the robotic arm. The program for implementing the control method of the robotic arm is executed by the processor to implement the steps of the control method of the robotic arm as described above.

[0037] The present application provides a control method for a robotic arm. The present application first performs time alignment on the edge gateway for the sensor data collected by the image acquisition module, the radio frequency identification tag, and the environmental sensor to obtain time synchronization data; performs noise reduction processing on the time synchronization data based on the modal type to obtain multimodal structured data; fuses the multimodal structured data according to the dynamic confidence determined by the ambient light data to obtain cargo features; inputs the cargo features into a sorting engine, combines the cargo priority and the equipment operation data to generate a sorting strategy; and adjusts the end position of the robotic arm and the torque parameters of the gripper based on the sorting strategy. This solves the technical problem in the related art that the large differences in multimodal data formats lead to low information integration efficiency, which in turn leads to low sorting efficiency. This achieves the technical effect of integrating multimodal data to improve sorting efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0040] Figure 1 A flowchart of the first embodiment of the control method of the robotic arm of the present application is provided;

[0041] Figure 2 This is a flowchart of steps S21-S23 in the third embodiment of the method for controlling a robotic arm of the present application;

[0042] Figure 3 This is a schematic diagram of the hardware structure of the sorting equipment involved in this application.

[0043] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0044] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0045] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0046] Currently, in the after-sales service field of smart devices, sorting devices are typically installed to enable sequential sorting and delivery of goods. Related technologies further utilize AI cameras, RFID tags, temperature and humidity sensors, and weight sensors to collect real-time cargo information and environmental data. This technology leverages AI visual recognition, the Internet of Things (IoT), and automated equipment integration to improve the accuracy and efficiency of warehouse operations. However, the multimodal data formats used in these technologies vary significantly, resulting in inefficient information integration and, consequently, low sorting efficiency.

[0047] This application achieves the technical effect of integrating multimodal data to improve sorting efficiency by integrating multimodal structured data.

[0048] It should be noted that the execution entity of this embodiment can be a storage and sorting system, or a computing service device with data processing, network communication, and program execution capabilities, such as a tablet computer, personal computer, or mobile phone, or a sorting device capable of performing the aforementioned functions, and this embodiment does not specifically limit this. The following uses the sorting device as an example to illustrate this embodiment and the following embodiments.

[0049] Based on this, the first embodiment of the present application proposes a control method for a robotic arm, please refer to Figure 1 , the control method of the robotic arm includes steps S10 to S50:

[0050] In step S10 , the sensor data collected by the image acquisition module, the radio frequency identification tag, and the environmental sensor are time-aligned at the edge gateway to obtain time-synchronized data.

[0051] In this embodiment, the image acquisition module is a high-definition camera array used to photograph goods in industrial scenarios. The RFID tag is a UHF electronic tag attached to the goods. The environmental sensors include devices that monitor temperature, humidity, and light. The edge gateway is a small computing device deployed on-site in the warehouse, responsible for receiving and processing the raw data from each sensor. Time alignment resolves the time asynchrony caused by different sampling frequencies of different sensors (for example, 30 frames per second for cameras and 10 frames per second for RFID), mapping all data to the same timeline.

[0052] As an optional implementation, the edge gateway first adds a hardware-level timestamp (accurate to microseconds) to each sensor data point and then uses the LAMP algorithm (an industrial time synchronization protocol) to calibrate the clock offset between each sensor and the gateway. For RFID and environmental sensor data with lower sampling frequencies, cubic spline interpolation is used to resample to the same 30Hz frequency as the camera, ensuring that multi-source data at the same time point can be analyzed accordingly. For example, if the camera captures an image at t=1000ms and the RFID tag reads the tag at t=1030ms, the gateway will interpolate the RFID data to the time point t=1000ms to align it with the image data.

[0053] Step S20: performing noise reduction processing on the time synchronization data based on the modality type to obtain multimodal structured data.

[0054] In this example, modality refers to the source of the data, which can be categorized as visual (camera image), electromagnetic (RFID signal), and physical (temperature, humidity, and light). Noise reduction is performed based on the characteristics of different modal data to remove interference. Multimodal structured data is the process of organizing processed data from various modalities into a unified format (such as JSON), containing clear fields such as cargo location, tag ID, and environmental parameters.

[0055] As an optional implementation, for the visual modality, a non-local mean filter is applied to the camera image to eliminate salt and pepper noise caused by light fluctuations during shooting. The bounding box coordinates and dimensions of the goods are then extracted using the YOLOv5 object detection model. For the electromagnetic modality, a Kalman filter is used on the RFID signal strength sequence to filter out signal jumps caused by reflections from metal shelves, retaining stable tag IDs and signal strength values. For the physical modality, a sliding median filter (window size 5) is used on the temperature and humidity data to eliminate outliers caused by transient sensor failures. Ultimately, the three types of processed data are integrated into structured data containing fields such as "location: (x, y), tag ID: P20250518001, temperature: 25°C, and light: 800 lux."

[0056] Step S30 : According to the dynamic confidence determined by the ambient light data, the multimodal structured data is integrated to obtain cargo features.

[0057] In this example, ambient light data refers to the ambient brightness value (in lux) collected by the light sensor. Dynamic confidence is a coefficient (ranging from 0.3 to 1.0) that adjusts the reliability of visual data based on light intensity. Cargo features are information that, after integration, comprehensively describes the cargo's attributes (such as type, size, and target area).

[0058] As an optional implementation, the edge gateway calculates the confidence level of the visual data under the current lighting conditions using a predefined illumination-confidence mapping function (for example, for every 100 lux increase in illumination, the confidence level increases by 0.1, with a minimum of 0.3). For example, at 800 lux, the confidence level is 0.6, with the visual data weighted at 60% and the RFID and environmental data weighted at 40%. The final product characteristics are calculated through a weighted summation: characteristic value = 0.6 × visual characteristics (e.g., dimensions 30 × 20 × 15 cm) + 0.4 × (RFID characteristics (tag ID) + environmental characteristics (suitable temperature)). This comprehensively determines that the product is a mobile phone packaging box and the target sorting area is the "electronics area."

[0059] Step S40: Input the cargo characteristics into the sorting engine, combine the cargo priority and equipment operation data, and generate a sorting strategy.

[0060] In this embodiment, the sorting engine is a path planning system deployed on an edge gateway or in the cloud. The order priority is determined by the order type (e.g., expedited orders have priority 5, standard orders have priority 1). Equipment operating data includes the current load on the robotic arm and conveyor speed. A sorting strategy is a sequence of instructions that includes a target location, a path, and an execution time.

[0061] As an optional implementation, the sorting engine uses a multi-objective optimization model with the shortest path, minimum latency, and lowest equipment load as optimization goals. It then uses the A* algorithm to search for the optimal path, taking into account item priorities (such as expedited orders) and equipment constraints (such as robot arm joint angle limits). For example, if a specific item is detected as an expedited order and robot arm A is at a load factor of 70% (which is below the upper limit), the engine prioritizes robot arm A's grasping path, avoiding other high-load equipment and generating a strategy: "Robot arm A → conveyor belt → sorting port C, estimated arrival time: 15 seconds."

[0062] Step S50: adjusting the end speed of the robot arm and the torque parameters of the gripper based on the sorting strategy.

[0063] In this example, the end-of-arm velocity is the linear velocity of the actuator (in m / s), and the gripper torque parameter is the force applied during the grip (in N·m). This adjustment process requires dynamic optimization based on the cargo's properties to ensure stable gripping without damaging the cargo.

[0064] As an optional implementation, the robotic arm controller first determines the base torque based on the cargo weight (obtained via RFID tags) by looking up a table (e.g., 12.5 N·m for a 0.5 kg cargo). It then adjusts the final torque based on the visually identified surface material (e.g., glass requires a 20% reduction in torque, cardboard requires a 20% increase). Simultaneously, the gripper motor's current feedback is monitored in real time. If current fluctuations exceed 15%, indicating slippage, the torque is automatically increased by 20%. For example, when grasping a 0.5 kg mobile phone box (made of cardboard), the base torque is 12.5 N·m, which is corrected to 15 N·m. The current remains stable during the grasping process, requiring no additional adjustment.

[0065] For example, an industrial camera captures an image of a shipment, an RFID tag reads its ID "P20250518001," and a light sensor records an ambient light level of 800 lux. After receiving this data, the edge gateway first timestamps the camera image (t=1000ms), RFID signal (t=1030ms), and light data (t=1010ms), then uses the LAMP algorithm to align the two to the timeline of t=1000ms. Next, the image is denoised to identify the shipment's dimensions as 30×20×15cm. The RFID signal is filtered to confirm the tag's validity, and the light data is filtered to remove outliers, retaining 800 lux. Based on the light intensity, a visual confidence level of 0.6 is calculated. By integrating visual and RFID features, the shipment is identified as a mobile phone packaging box and the target sorting area is the "electronics area." The sorting engine, taking into account the order priority (expedited) and the load factor of Robot A (70%), plans the optimal route: Robot A → conveyor belt → sorting port C, with an estimated arrival time of 15 seconds. The robot controller sets the gripper torque to 15 N·m and the terminal speed to 0.8 m / s based on the cargo weight (0.5 kg) and material (cardboard). The current remains stable during the gripping process, allowing for smooth sorting.

[0066] This application first aligns the sensor data collected by the image acquisition module, radio frequency identification tag and environmental sensor at the edge gateway to obtain time synchronization data; performs noise reduction processing on the time synchronization data based on the modal type to obtain multimodal structured data; fuses the multimodal structured data according to the dynamic confidence determined by the ambient light data to obtain cargo features; inputs the cargo features into the sorting engine, combines the cargo priority and equipment operation data to generate a sorting strategy; adjusts the end position of the robot arm and the torque parameters of the gripper based on the sorting strategy. This solves the technical problem in the related art that the large differences in multimodal data formats lead to low information integration efficiency, which in turn leads to low sorting efficiency. This achieves the technical effect of integrating multimodal data to improve sorting efficiency.

[0067] Based on any embodiment, in the second embodiment of the present application, step S10 includes:

[0068] Step S11 : receiving the image stream collected by the image acquisition module, the RFID data collected by the RFID tag, and the environmental data collected by the environmental sensor.

[0069] In this embodiment, the image stream is a sequence of cargo images continuously captured by an industrial camera at a fixed frame rate (e.g., 30 frames per second). The radio frequency identification data is the ID and signal strength value read from the cargo tag by a UHF RFID reader (e.g., 10 times per second). The environmental data is the environmental parameters collected in real time by the temperature and humidity sensor and the light sensor (e.g., 10 times per second).

[0070] As an optional implementation, the edge gateway receives the image stream transmitted by the camera via industrial Ethernet (such as Profinet), reads the radio frequency identification data from the RFID reader via the RS485 interface, and obtains the temperature, humidity, and light values from the environmental sensors via the Modbus protocol. All data is automatically timestamped by the edge gateway's local timestamp (accurate to the millisecond) upon receipt. For example, the first frame of the image stream has a timestamp of t=1000ms, and the second frame has a timestamp of t=1033ms; the first RFID data read has a timestamp of t=1000ms, and the second has a timestamp of t=1100ms; the first environmental data collection has a timestamp of t=1000ms, and the second has a timestamp of t=1100ms.

[0071] Step S12 : performing interpolation on the image stream based on a cubic spline according to the frame rate of the image stream and the acquisition frequency of the RFID data, so as to align the image stream and the RFID data.

[0072] In this embodiment, the frame rate of the image stream is the number of frames captured by the camera per second (e.g., 30 fps corresponds to one frame every 33 ms), and the frequency of RFID data acquisition is the number of times the RFID reader reads data per second (e.g., 10 Hz corresponds to one frame every 100 ms). Cubic spline interpolation fits the time-pixel relationship of the image stream using a piecewise cubic polynomial to generate additional interpolated frames to match the time points of the RFID data.

[0073] As an optional implementation, the time interval between the image stream and the RFID data is first calculated: the time interval for the 30fps image stream is 33ms (1000ms / 30), and the time interval for the 10Hz RFID data is 100ms (1000ms / 10). To align the two, cubic spline interpolation is performed on the image stream, targeting the time points of the RFID data (t=1000ms, 1100ms, 1200ms, etc.). For example, at RFID time point t=1100ms, the original image stream frames are t=1066ms (frame 3) and t=1100ms (frame 4). Cubic spline interpolation is used to generate an interpolated frame at t=1100ms, ensuring that both image data and RFID data are present at this time point.

[0074] Step S13 : marking the environmental data at each time stamp, and the image stream and the RFID data aligned at the time stamp as the time synchronization data.

[0075] In this embodiment, the timestamp is a specific time point of data collection (eg, t=1000ms, 1100ms), and the time synchronization data is a collection of image frames, RFID data, and environmental data associated with the same timestamp.

[0076] As an optional implementation, the edge gateway traverses the timestamps of all RFID data (such as t=1000ms, 1100ms, 1200ms), and for each timestamp, obtains the image frame after cubic spline interpolation at that time point (such as the interpolated frame at t=1100ms), and at the same time searches for the environmental data collected by the environmental sensor near the timestamp (such as temperature and humidity of 25°C and light of 800lux at t=1100ms), and packages the three into a time synchronization data unit containing "timestamp: 1100ms, image: interpolated frame, RFID: tag IDP20250518001, temperature and humidity: 25°C, light: 800lux".

[0077] For example, an industrial camera captures images of goods on the sorting line at 30 fps (one frame every 33 ms), an RFID reader reads the goods tags at 10 Hz (once every 100 ms), and an environmental sensor collects temperature, humidity, and light data 10 times per second. The edge gateway receives the camera image stream via Industrial Ethernet, acquires RFID data via RS485, and reads environmental data via the Modbus protocol. All data is locally timestamped. When the RFID reader reads the tag ID "F20250518001" at t=1100 ms, the image stream contains original images at t=1066 ms (frame 3) and t=1100 ms (frame 4). The gateway uses cubic spline interpolation to generate an interpolated image frame at t=1100 ms. Simultaneously, the environmental sensor captures temperature and humidity of 22°C and light of 750 lux at t=1100 ms. The gateway associates the interpolated image at t=1100 ms with the RFID and environmental data to form a time-synchronized data unit containing complete goods information at that point in time.

[0078] This embodiment solves the time asynchrony problem caused by different acquisition frequencies of image streams, RFID data, and environmental data by receiving and time-aligning multi-source data. It ensures that multimodal data with the same timestamp can be correlated and analyzed, providing a time-consistent input basis for subsequent cargo feature fusion, and avoiding problems such as image and label mismatches and incorrect associations between environmental parameters and cargo status caused by time misalignment.

[0079] Based on any of the above embodiments, in the third embodiment of the present application, refer to Figure 2 Step S20 includes:

[0080] Step S21: performing noise reduction processing on the time synchronization data based on the modality type of the time synchronization data.

[0081] In this embodiment, the modality type refers to the characteristics of the data source, which can be divided into three categories: visual (image stream), electromagnetic (RFID signal), and physical (environmental parameters). Noise reduction processing is a noise suppression algorithm for data with different characteristics.

[0082] As an optional implementation, for the visual modality, a non-local mean filter (window size 7×7, filter strength h=10) is applied to the image stream to eliminate salt and pepper noise caused by changes in illumination. For the electromagnetic modality, a Kalman filter (state transfer matrix A=1, observation matrix H=1) is used on the RFID signal strength sequence to remove signal jumps caused by reflections from metal shelves. For the physical modality, a sliding median filter (window size 5) is applied to the temperature and humidity data to eliminate abnormal pulses caused by instantaneous sensor jitter.

[0083] Step S22: determining abnormal data in the time synchronization data based on the rule engine and the equipment operation data in the warehouse.

[0084] In this embodiment, the rule engine is a business rule execution system based on the Drools framework. The equipment operation data includes real-time parameters such as conveyor belt speed and robot arm load rate. Abnormal data is data that does not conform to the preset rules or equipment status logic.

[0085] As an optional implementation, the rule engine loads the following rule sets: Image integrity rule: If the image grayscale standard deviation is less than 10 (indicating an image that is too dark or overexposed), it is marked as an anomaly; RFID signal rule: If the signal strength RSSI is less than -80dBm and the fluctuation amplitude is greater than 15dBm, it is considered a signal anomaly; Environmental parameter rule: If the temperature is greater than 40°C or the humidity is greater than 90%RH (outside the normal storage range), it is marked as an anomaly; Equipment association rule: If the conveyor speed is greater than 2m / s but the position of the goods in the image does not change, the image and equipment status are considered inconsistent. When the time synchronization data triggers any of these rules, the corresponding data is marked as an anomaly.

[0086] Step S23: filtering the abnormal data from the denoising result and encapsulating the result into the multimodal structured data.

[0087] In this embodiment, the denoising result is the output of each modal data after processing in step S20, the abnormal data is the data marked in step S30, and the multimodal structured data is the integrated data that complies with the format specification and has no abnormalities.

[0088] As an optional implementation, the noise-reduced data from each modality is processed according to the following process: Data filtering: Remove time-synchronized data units marked as abnormal by the rule engine (such as abnormal image frames, RFID signals, and environmental parameters); Format unification: Convert visual data into a JSON object containing "location coordinates (x, y), dimensions (w, h), and category ID", RFID data into "tag ID, RSSI value", and environmental data into "temperature, humidity, and light intensity"; Data association: Using timestamps as keys, encapsulate visual, RFID, and environmental data at the same time point into a nested JSON structure:

[0089] json{"timestamp": 1100,

[0090] "vision": {"position": [100, 200], "size": [30, 20], "class": "Mobile phone"},

[0091] "rfid": {"tag_id": "P20250518001", "rssi": -65},

[0092] "environment": {"temperature": 25, "humidity": 60, "light": 800}}.

[0093] For example, the time-synchronized data includes an image frame at t = 1100ms (showing the location of the cargo), an RFID signal (tag ID "P20250518001"), and environmental parameters (temperature 25°C, 800 lux). The image frame is filtered using a non-local means filter to remove noise, the RFID signal is filtered using a Kalman filter to stabilize the RSSI at -65dBm, and the environmental data is filtered using a median filter to verify that no anomalies are present. The rules engine verifies that all data meets pre-set rules (e.g., image grayscale standard deviation of 50, RSSI fluctuation <10dBm, and temperature within the normal range). It then encapsulates the visual, RFID, and environmental data at that point in time into multimodal structured data in JSON format for subsequent sorting decisions.

[0094] Optionally, the time synchronization data further includes gyroscope data and device operation data, the device operation data including the speed of the end of the robot arm and the opening and closing angle of the gripper. The step of performing noise reduction processing on the time synchronization data based on the modal type of the time synchronization data includes any one of the following:

[0095] Step S211 : Decomposing the gyroscope data by three-level wavelet decomposition to obtain a position signal with a preset accuracy.

[0096] Step S212 , dynamically compensating for random drift errors in the equipment operation data, removing singular values in combination with the Labuda criterion, and removing trend items by the least squares method.

[0097] Step S213: The image stream is subjected to adaptive histogram equalization.

[0098] In this embodiment, the time synchronization data is newly added with gyroscope data (used to detect the device posture) and device operation data (including the speed of the end of the robot arm and the opening and closing angle of the gripper), and noise reduction needs to be performed based on different modal characteristics.

[0099] As an optional implementation, the following processing is used: Gyroscope data processing: The raw gyroscope angular velocity data is decomposed into approximate coefficients (low-frequency) and detail coefficients (high-frequency) using the Daubechies4 wavelet. The approximate coefficients (representing the true motion trend) are retained, and the detail coefficients are subjected to soft threshold denoising (threshold λ = σ√(2lnN), where σ is the noise standard deviation). The wavelet coefficients are reconstructed to obtain a smoothed angular velocity signal, which is then integrated to convert into a position signal with an accuracy of ±0.5°. Equipment operation data processing: Dynamic compensation for random drift error: An ARIMA(1,1,1) model is established to predict velocity / angle drift and the compensation amount is adjusted in real time. The Labuda criterion is used to eliminate singular values: The data standard deviation σ is calculated. If a data point deviates from the mean by >3σ, it is considered a singular value and is eliminated. The least squares method is used to remove trend terms: A quadratic polynomial is fitted to the velocity / angle series, and the detrended signal is obtained by subtracting the fitted curve from the original data. Image stream processing: CLAHE is performed on the image frame, dividing the image into 8×8 sub-blocks. Histogram equalization is performed independently on each sub-block, limiting the contrast enhancement factor to 40. Bilinear interpolation is used to eliminate artifacts at sub-block boundaries and improve image details (such as the clarity of the edges of the goods).

[0100] This embodiment effectively suppresses noise interference from different sensor data types through targeted modal noise reduction. A rule engine, combined with device operating status, automatically identifies and filters abnormal data. Finally, multi-source data is integrated into a standardized JSON structure, providing uniformly formatted, reliable input data for subsequent cargo feature fusion, ensuring that sorting decisions are based on accurate and consistent information. Targeted noise reduction resolves high-frequency noise in gyroscope data, improving attitude detection accuracy. Singular values and trend errors in device operating data are eliminated, ensuring the stability of control commands. Image detail is enhanced, improving visual recognition accuracy.

[0101] Optionally, the abnormal data includes format error data and out-of-limit data, and the step of determining the abnormal data in the time synchronization data according to the rule engine and the equipment operation data in the warehouse includes:

[0102] Step S221: determining format error data in the time synchronization data based on the rule engine.

[0103] In this embodiment, the rule engine is a business rule execution system based on the Drools framework, and format error data is data that does not conform to the preset data structure or field constraints.

[0104] As an optional implementation, the rule engine loads the following format verification rules: Image data rule: If the image size is not 1920×1080 pixels or the grayscale value is outside the range of 0-255, it is determined to be format incorrect; RFID data rule: If the tag ID length is ≤ 12 bits or the RSSI value is outside the range of [-100, -20] dBm, it is determined to be format incorrect; Environmental data rule: If the temperature value is outside the range of [-20, 60]°C or the humidity value is outside the range of [0, 100]%RH, it is determined to be format incorrect; Gyroscope data rule: If the angular velocity value is outside the range of [-300, 300]° / s, it is determined to be format incorrect. When time synchronization data triggers any of these rules, the corresponding data is marked as format incorrect.

[0105] Step S222: construct an isolation tree based on the equipment operation data, cargo weight, and each of the time synchronization data.

[0106] In this embodiment, the equipment operation data includes the speed of the end of the robot arm and the opening and closing angle of the gripper. The cargo weight is obtained through a pressure sensor. The isolation tree is an anomaly detection model built based on the IsolationForest algorithm.

[0107] As an optional implementation, the following features are input into the isolation tree model: robot arm end speed (m / s); gripper opening and closing angle (°); cargo weight (kg); ambient temperature (°C); and RFID signal strength (RSSI, dBm). A feature dimension (such as speed) is randomly selected, and a split point is randomly selected between the maximum and minimum values of that dimension. The data space is recursively partitioned until each data point is isolated or the maximum depth of the tree (default 8 levels) is reached.

[0108] Step S223 : determining out-of-limit data based on the path length and anomaly score of each data point in the isolated tree.

[0109] In this embodiment, the path length is the number of edges from the root node to the leaf node of a data point in the isolation tree, and the anomaly score is an outlier degree indicator (ranging from 0 to 1) calculated based on the path length.

[0110] As an optional implementation, the average path length E(h(x)) of each data point in all isolated trees is calculated and converted into an anomaly score: s(x,n)=2^{-E(h(x)) / c(n)}.

[0111] Where c(n) is the average path length of the tree, and n is the number of samples. When the score s(x,n) > 0.7, the data point is considered an outlier. For example, if the robot arm speed is far outside the normal range (e.g., 5 m / s, where normal is ≤ 1 m / s), the anomaly score might reach 0.95; if the gripper angle does not match the cargo weight (e.g., the angle is too small when gripping a 5 kg cargo), the anomaly score might reach 0.82.

[0112] For example, in smart warehousing, time-synchronized data includes: robotic arm speed 0.8 m / s, gripper angle 60°, cargo weight 3 kg, temperature 25°C, and RSSI -65 dBm. The rule engine detects that all data formats are normal (for example, a temperature of 25°C is within the range [-20, 60]°C). The isolation tree model calculates an anomaly score of 0.2 for this data point (well below the threshold of 0.7), identifying it as normal data. If abnormal data occurs at a certain moment (for example, a sudden increase in robotic arm speed to 3 m / s), the rule engine still determines the format is correct, but the isolation tree model calculates an anomaly score of 0.9, identifying it as out-of-limit data.

[0113] This embodiment uses a rule engine to quickly identify formatted data, and uses the isolation forest algorithm to build an anomaly detection model based on the equipment operating status and environmental parameters. It can automatically identify out-of-limit data that exceeds the normal range, avoid erroneous decisions caused by sensor failure or temporary interference, and provide reliable data quality assurance for the stable operation of the sorting equipment.

[0114] Based on any of the above embodiments, in the fourth embodiment of the present application, before step S30, the following steps are included:

[0115] Step A10: extracting the visual feature vector corresponding to the image stream through a feature extraction module.

[0116] In this embodiment, the feature extraction module uses industrial resolution data input. The first layer of the feature extraction module uses a large 5×5 convolution kernel, and subsequent layers use smaller 3×3 convolution kernels. Each layer fuses a 128-channel conv3_x, a 256-channel cnv4_x, and a 512-channel conv5_x to output a 256-dimensional visual feature vector. The feature extraction module is a convolutional neural network based on an improved ResNet architecture. Industrial resolution refers to an input image size of 1280×720 pixels. The first layer uses large convolution kernels to capture global image features, while subsequent smaller convolution kernels extract details. A squeeze-and-excitation (SE) module is used to enhance key channel feature responses.

[0117] As an optional implementation, the feature extraction module is implemented according to the following structure: input layer: receives RGB images with industrial resolution (1280×720); first convolution layer: uses a 5×5 large convolution kernel (step 2, padding 2) to reduce the image dimension to 640×360 and extract global edge and texture features; middle layer: uses a 3×3 small convolution kernel (step 1, padding 1) to construct a residual block, and generates in sequence: conv3_x: 128-channel feature map (size 320×180); conv4_x: 256-channel feature map (size 160×90); conv5_x: 512-channel feature map (size 80×45); feature fusion: conv3 _x, conv4_x, and conv5_x are upsampled to the same size (160×90) through bilinear interpolation; after splicing according to the channel dimension, they are compressed to 256 dimensions through 1×1 convolution to generate the initial feature vector; attention enhancement: the SE module is inserted after conv4_x to perform global average pooling on the 256-channel feature map to generate a 256-dimensional channel descriptor; a channel weight vector is generated through a two-layer fully connected network (W1: 256→64, W2: 64→256) and an activation function (δ is ReLU, σ is Sigmoid); the channel weight is multiplied by the original feature map channel by channel to enhance the response of key channels (such as cargo edges and texture features).

[0118] For example, an industrial image (1280×720) containing a mobile phone box is input, and the feature extraction module processing flow is as follows: the first layer 5×5 convolution captures the overall outline of the box and generates a 640×360 feature map; the middle layer 3×3 convolution extracts details such as the surface texture and printed text of the box, and conv4_x generates a 256-channel feature map; the SE module performs global average pooling on the conv4_x feature map to obtain a 256-dimensional channel descriptor; the fully connected network calculates the channel weight, such as the "box edge" channel weight is increased to 0.9, and the "background noise" channel weight is reduced to 0.2; the weighted feature map is fused with conv3_x and conv5_x to generate a 256-dimensional visual feature vector, which accurately characterizes the shape, size and surface features of the mobile phone box.

[0119] This embodiment balances the ability to capture global features and extract details by combining large convolution kernels with small convolution kernels; multi-scale feature fusion enhances the richness of feature expression; the SE module adaptively adjusts channel weights to highlight key features and suppress noise, making the generated visual feature vector more robust to interference such as lighting changes and slight occlusion, providing high-quality visual input for subsequent multimodal fusion.

[0120] Based on any of the above embodiments, in the fifth embodiment of the present application, step S30 includes:

[0121] Step S31, determining the dynamic confidence level by looking up a table based on the ambient light data;

[0122] Step S32: fusing the visual feature vector and the radio frequency identification data according to the dynamic confidence to obtain the cargo feature.

[0123] In this embodiment, the ambient light data is the real-time brightness value (unit: lux) collected by the light sensor, and the dynamic confidence is the coefficient (range: 0.3-1.0) that adjusts the credibility of the visual feature based on the light intensity. As an optional implementation, a mapping table of light intensity and confidence is pre-established: light intensity interval (lux) dynamic confidence. 0-200 corresponds to 0.3. 201-500 corresponds to 0.5. 501-1000 corresponds to 0.7. 1001-2000 corresponds to 0.9. >2000 corresponds to 0.7. When the ambient light data is received, the corresponding interval is quickly located through binary search to determine the dynamic confidence. For example, when the light intensity is 800 lux, the confidence obtained by looking up the table is 0.7.

[0124] In this embodiment, the visual feature vector is an extracted 256-dimensional vector, the RFID data includes the tag ID and signal strength (RSSI), and the cargo feature is a fused comprehensive representation (such as cargo type, location, and priority).

[0125] As an optional implementation, a weighted fusion strategy is adopted: Visual feature processing: The 256-dimensional visual feature vector v is input into the fully connected layer and mapped into the cargo attribute vector v' (for example, [0.8, 0.1, 0.1] represents "mobile phone: 80%, tablet: 10%, other: 10%"); RFID feature processing: The RSSI value is converted into the existence probability p (range 0-1) using the function f(RSSI) = 1 / {1+e^{-k(RSSI-b)}} (k=0.2, b=-70); Weighted fusion: cargo feature = a*v'+(1-a)*[p,0,0].

[0126] Where a is the dynamic confidence level. For example, when the illumination is 800 lux (a = 0.7), if the visual recognition probability of a mobile phone is 0.8 and the RFID presence probability is 0.9, the fused cargo features are: [0.83, 0.07, 0.07].

[0127] This embodiment dynamically adjusts the visual feature weights by light intensity, solving the robustness problem of traditional fixed-weight fusion methods in low-light or strong-light environments.

[0128] Based on any of the above embodiments, in the sixth embodiment of the present application, step S40 includes:

[0129] Step S41 : The cargo characteristics, the cargo priority parsed from the RFID data, and the device operation data are combined into a state vector.

[0130] In this embodiment, the cargo feature is the fused feature vector output in step S40 (e.g., [0.718, 0.18, 0.092]), the cargo priority is an integer value parsed from the RFID tag data (e.g., levels 1-5), and the equipment operation data includes the current speed of the robotic arm (m / s), the gripper angle (°), etc.

[0131] As an optional implementation, the above data is concatenated into a 256-dimensional state vector in the following format: the first three dimensions: cargo characteristics (e.g., [0.718, 0.18, 0.092]); the fourth dimension: cargo priority (normalized to [0, 1], e.g., priority 3 → 0.6); and the fifth to second dimensions: equipment operation data for each device (robotic arm speed and gripper angle, normalized to [0, 1]). If the data has fewer than 256 dimensions, zeros are added.

[0132] Step S42: Input the state vector into the sorting engine, and use the output of the sorting engine as the sorting strategy. The sorting engine includes a 256-dimensional input layer, a two-layer hidden layer with 128 units for processing time series data, and an output layer with 51 discretized action branches, where the action branches include the robot arm end speed and the gripper opening and closing angle.

[0133] In this embodiment, the sorting engine is a reinforcement learning model based on LSTM. The input layer receives a 256-dimensional state vector, processes the timing dependency through two layers of 128-unit LSTM, and the output layer is discretized into 51 action branches (corresponding to different combinations of robot arm speed and gripper angle).

[0134] As an optional implementation, the sorting engine's network architecture is as follows: Input layer: 256-dimensional fully connected layer, receiving the state vector; Hidden layer: 1st LSTM layer: 128 units, returning a sequence (processing time series data); 2nd LSTM layer: 128 units, returning the final state; Output layer: 51 fully connected branches, each outputting a scalar value (Q value); Action branches 1-25: End-of-arm velocity (0.1-2.5 m / s, 0.1 m / s step); Action branches 26-51: Gripper opening and closing angle (30°-180°, 6° step). The optimal action is selected using an ε-greedy strategy: a^*=argmaxQ(s,a). Here, s is the state vector, A is the action space, and Q(s,a) is the action-value function.

[0135] For example, the state vector at a certain moment is: [0.718, 0.18, 0.092, 0.6, 0.4, 0.5, 0.2, 0.7, ..., 0.2], which indicates: the probability of the item being a mobile phone is 71.8%, the priority is 3, the current speed of the robot arm is 0.8 m / s (normalized value 0.4), the gripper angle is 90° (normalized value 0.5), etc. After processing by the sorting engine, the Q values of each action branch in the output layer are as follows: speed branch: Q(0.8 m / s) = 0.92, Q(1.0 m / s) = 0.95, Q(1.2 m / s) = 0.88; angle branch: Q(90°) = 0.90, Q(108°) = 0.96, Q(120°) = 0.85. The ε-greedy strategy is used to select the combination with the highest Q value: robot arm speed 1.0m / s, gripper angle 108°, and the generated sorting strategy is: {sorting strategy}=[1.0m / s,108 degrees].

[0136] For example, sensor data collected by various sensors is received and then time-series synchronized at the edge gateway. A noise reduction method is selected based on the modality of the sensor data. The processed data is then sent to the cloud. The cloud uses a rules engine and anomaly detection to filter out abnormal data, and an intelligent decision engine generates a sorting strategy. Based on this sorting strategy, the robot arm's end velocity and gripper torque parameters are analyzed, and path planning is performed for each robot arm.

[0137] For example, this embodiment builds a closed-loop "perception-computation-decision-execution" system around the automated sorting needs of warehousing scenarios. Through the in-depth cooperation of four modules: multimodal sensor collaboration, edge computing optimization, dynamic data fusion, and intelligent decision-making, it achieves full automation of the entire process from information collection to precise sorting of goods. By building a multi-dimensional perception network, it breaks through the limitations of traditional single-modal data collection, optimizes sensor signal synchronization and anti-interference capabilities, and improves recognition accuracy in complex scenarios (such as light changes). Edge intelligent gateways are deployed, integrating MEMS gyroscopes and noise reduction algorithms to achieve local data preprocessing and millisecond-level response. Combined with a rule engine, the Transformer algorithm fusion architecture is adopted to simultaneously optimize sorting path planning and exception handling mechanisms.

[0138] Multimodal sensor integration: Image and data acquisition is performed using AI cameras, RFID tags, and environmental sensors. A multimodal collaborative model-based bionic perception system integrates multiple sensing modalities (such as touch and chemical signals) to achieve deep integration of environmental perception and intelligent decision-making. Device collaboration and edge computing nodes: Gateway devices are deployed at the warehouse edge to preprocess local data (such as filtering noise and standardizing the format) before uploading it to the cloud via 5G / IoT protocols. Specifically, the ARMA model is used to dynamically compensate for random drift errors, the Labuda criterion (3σ) is used to eliminate singular values, and the least squares method is used to remove trend terms to reduce raw data noise. A real-time noise reduction network based on a lightweight CNN is introduced, with only 180 parameters, achieving 30fps inference in embedded devices. The gyroscope signal is decomposed into a three-level structure using the Daubechies wavelet basis. High-frequency noise components are extracted using the modulus maximum method. A soft threshold (threshold = 0.5σ) is used to compress the noise subband coefficients, retaining the valid signal with an accuracy of ±3mm. The edge coefficient is optimized by combining Contourlet transform, the coarse-grained noise is weakened by the shrinkage factor α (0.8-0.9), and the weak edges are enhanced by the protrusion factor β (1.2-1.5), achieving a 40% improvement in the signal-to-noise ratio.

[0139] Invalid data is eliminated through rule engines such as regular expressions and anomaly detection algorithms such as isolation forests. This is achieved by optimizing the weights of multi-source data in IoT resource scheduling, for example by dynamically adjusting the confidence level of AI camera and RFID data based on cargo priority. Specifically, data elimination is achieved by constructing multiple isolation trees to isolate data points. The length of the path through which a data point is isolated in the tree determines whether it is an outlier (invalid data). Data partitioning and tree construction: For various types of data collected in the warehouse (such as cargo weight, inventory quantity, and equipment operating parameters), a data sample is used as input. A feature is randomly selected, and a split point is randomly chosen within the feature's value range to divide the data space into two parts. This process is repeated to construct multiple isolation trees. Path length calculation: For each data point, the path length from the root node to the leaf node in each isolation tree is calculated. Because their feature values differ from those of most normal data points, outliers are often isolated to leaf nodes after fewer splits, resulting in shorter paths. Normal data points, on the other hand, typically reach leaf nodes after more splits and have longer paths. Outlier Identification and Removal: Calculate an anomaly score for each data point based on the path lengths of all data points in the isolated tree. Generally speaking, the higher the anomaly score, the more likely the data point is an outlier (invalid data). Set an appropriate threshold. When a data point's anomaly score exceeds this threshold, it is considered invalid and removed.

[0140] The implementation of a CNN-based, reinforcement learning-based sorting engine requires a deep fusion of visual perception and dynamic decision-making technologies. Data fusion results trigger automated equipment actions, linking robotic arms and intelligent dispatching robots to automatically sort goods. State-space modeling defines the state vector St = [fcnn, P, v, θ], which contains: the CNN feature vector fcnn∈R256 (fully connected layer output); the item priority P∈{1−5}; the robotic arm end velocity v∈[0, 2.3 m / s]; and the gripper opening and closing angle θ∈[0°, 180°]. Visual data (30 fps) and RFID data (10 Hz) are aligned using cubic spline interpolation. Online homography calibration is performed: the transformation matrix H is updated every 500 ms, optimizing the objective function:

[0141] .

[0142] where X cam is the feature vector generated by the data of the image stream, X RFID is the feature vector generated by the data collected by the RFID tag. N is the dimension. F is the standard mathematical symbol for the Frobenius norm.

[0143] This embodiment integrates multi-source information through state vector splicing, uses the LSTM network to effectively handle the timing dependency problem in warehousing and sorting, and the discretized action space makes the control instructions more precise.

[0144] The present application provides a sorting device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the control method of the robotic arm in the above-mentioned embodiment one.

[0145] Reference below Figure 3 , which shows a schematic diagram of the structure of a sorting device suitable for implementing the embodiments of the present application. The sorting device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 3 The sorting device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.

[0146] like Figure 3As shown, the sorting device may include a processing device 1001 (e.g., a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1003 into a random access memory (RAM) 1004. RAM 1004 also stores various programs and data required for the operation of the sorting device. The processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008, such as a liquid crystal display (LCD), speaker, vibrator, etc.; storage device 1003, such as a magnetic tape or hard disk; and communication devices 1009. The communication device 1009 can allow the sorting device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a sorting device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.

[0147] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0148] The sorting device provided in this application, utilizing the robotic arm control method described in the aforementioned embodiment, can address the technical issues of large differences in multimodal data formats, resulting in inefficient information integration and, consequently, low sorting efficiency. Compared to the prior art, the sorting device provided in this application achieves the same beneficial effects as the sorting device provided in the aforementioned embodiment. Other technical features of this sorting device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0149] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0150] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0151] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, and the computer-readable program instructions are used to execute the control method of the robotic arm in the above-mentioned embodiment.

[0152] The computer-readable storage medium provided herein may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including, but not limited to, wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0153] The computer-readable storage medium may be included in the sorting device, or may exist independently without being assembled into the sorting device.

[0154] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the sorting device, the sorting device: performs time alignment on the sensor data collected by the image acquisition module, the radio frequency identification tag, and the environmental sensor at the edge gateway to obtain time-synchronized data; performs noise reduction processing on the time-synchronized data based on the modal type to obtain multimodal structured data; obtains cargo characteristics by fusing the multimodal structured data based on the dynamic confidence determined by the ambient light data; inputs the cargo characteristics into the sorting engine, combines the cargo priority and the equipment operation data to generate a sorting strategy; and adjusts the end speed of the robot arm and the torque parameters of the gripper based on the sorting strategy.

[0155] Computer program code for performing the operations of the present application may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0156] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0157] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0158] The computer-readable storage medium provided in this application is a computer-readable storage medium storing computer-readable program instructions (i.e., a computer program) for executing the aforementioned robotic arm control method. This computer-readable storage medium can address the technical issues of large differences in multimodal data formats, resulting in inefficient information integration and, consequently, low sorting efficiency. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are similar to those of the robotic arm control method provided in the aforementioned embodiments and are not further elaborated here.

[0159] An embodiment of the present application provides a computer program product, including a computer program, which implements the steps of the above-mentioned method for controlling a robotic arm when executed by a processor.

[0160] The computer program product provided in this application can address the technical problem of inefficient information integration and, consequently, low sorting efficiency caused by large differences in multimodal data formats. Compared to the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are similar to those of the robotic arm control method provided in the aforementioned embodiments, and are not further elaborated here.

[0161] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the present application.

Claims

1. A method for controlling a robotic arm, characterized in that: The control method of the robotic arm includes: The sensor data collected by the image acquisition module, RFID tags, and environmental sensors are time-aligned at the edge gateway to obtain time-synchronized data; performing noise reduction processing on the time synchronization data based on the modality type of the time synchronization data; Determining abnormal data in the time synchronization data based on a rule engine and equipment operation data in the warehouse, wherein the time synchronization data also includes gyroscope data and equipment operation data, and the equipment operation data includes the speed of the end of the robot arm and the opening and closing angle of the gripper; Filtering the abnormal data from the denoising result and encapsulating it into multimodal structured data; fusing the multimodal structured data to obtain cargo features based on a dynamic confidence level determined by the ambient light data; Input the cargo characteristics into the sorting engine, combine the cargo priority and equipment operation data, and generate a sorting strategy; Adjusting the end speed of the robot arm and the torque parameters of the gripper based on the sorting strategy; The step of performing noise reduction processing on the time synchronization data based on the modality type of the time synchronization data includes any one of the following: Decomposing the gyroscope data by three-level wavelet decomposition to obtain a position signal with a preset accuracy; The equipment operation data is subjected to dynamic compensation for random drift errors, the Labuda criterion is combined to eliminate singular values, and the trend term is removed by the least square method; Pass the image stream through adaptive histogram equalization.

2. The method for controlling a robotic arm according to claim 1, wherein: The step of time-aligning the sensor data collected by the image acquisition module, the radio frequency identification tag, and the environmental sensor at the edge gateway to obtain time-synchronized data includes: receiving the image stream captured by the image acquisition module, the RFID data captured by the RFID tag, and the environmental data captured by the environmental sensor; interpolating the image stream based on a cubic spline according to a frame rate of the image stream and an acquisition frequency of the RFID data to align the image stream and the RFID data; Based on the environmental data at each time stamp, and the image stream and the RFID data aligned at the time stamp, it is marked as the time synchronization data.

3. The control method of the robotic arm according to claim 1, wherein: The abnormal data includes format error data and out-of-limit data. The step of determining the abnormal data in the time synchronization data based on the rule engine and the equipment operation data in the warehouse includes: determining, based on the rules engine, formatted data in the time synchronization data; Building an isolation tree based on the equipment operation data, cargo weight, and each of the time synchronization data; Out-of-limit data is determined based on the path length and anomaly score of each data point in the isolation tree.

4. The method for controlling a robotic arm according to claim 1, wherein: Before the step of fusing the multimodal structured data with the dynamic confidence determined based on the ambient light data to obtain cargo features, the method includes: The visual feature vector corresponding to the image stream is extracted through the feature extraction module; wherein, the data input format of the feature extraction module is industrial resolution; the first layer of the feature extraction module is a 5×5 large convolution kernel, and the subsequent layers use a 3×3 small convolution kernel; each layer is connected and fused with a 128-channel conv3_x, a 256-channel cnv4_x, and a 512-channel conv5_x, and then the fusion output is a 256-dimensional visual feature vector.

5. The method for controlling a robotic arm according to claim 1, wherein: The step of fusing the multimodal structured data with the dynamic confidence determined based on the ambient light data to obtain cargo characteristics includes: Determining the dynamic confidence level by looking up a table based on the ambient light data; The cargo features are obtained by fusing the visual feature vector and the radio frequency identification data according to the dynamic confidence.

6. The method for controlling a robotic arm according to claim 1, wherein: The step of inputting the cargo characteristics into the sorting engine and generating a sorting strategy based on cargo priority and equipment operation data includes: splicing the cargo characteristics, the cargo priority parsed from the RFID data, and the device operation data into a state vector; The state vector is input into the sorting engine, and the output of the sorting engine is used as the sorting strategy. The sorting engine includes a 256-dimensional input layer, a two-layer hidden layer with 128 units for processing time series data, and an output layer that discretizes 51 action branches, where the action branches include the end velocity of the robot arm and the opening and closing angle of the gripper.

7. A sorting device, characterized in that: The sorting device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for controlling the robot arm according to any one of claims 1 to 6.

8. A storage medium, characterized in that: The storage medium is a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for controlling the robotic arm according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Mechanical arm control method and system based on multi-mode driving and storage medium

    CN118752495A

  • Navigation method of multi-mode mobile sorting robot and sorting system

    CN119635669A

  • Intelligent robot control method based on multi-sensor cooperation

    CN119990181A