Image-text warehouse-out sorting management system

The QR code information and appearance images of graphic and text products are obtained by scanning the device, and semantic analysis is carried out in combination with deep learning algorithms, which solves the problem that the information in the sorting of graphic and text products is inconsistent with the actual object, realizing accurate identification and efficient sorting.

CN120494656AInactive Publication Date: 2025-08-15ZHUJI HEWU DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510612141.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-15
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When handling graphic and text products, the existing out-of-stock sorting system relies on the accuracy of early data entry, which leads to the problem that the product information does not match the physical objects, which is difficult to detect and correct in a timely manner, resulting in sorting errors.

Method used

The QR code information of graphic and text products is obtained by scanning the device, and the appearance images of their graphic and text are collected. The semantic analysis is performed using deep learning artificial intelligence algorithms to verify the accuracy of product information, and then the information is consistent with the actual object and then automatically transmitted and sorted.

Benefits of technology

It effectively reduces the error rate, realizes accurate identification and efficient sorting of graphic and text products, and ensures the consistency between product information and physical objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494656A_ABST
    Figure CN120494656A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition, and particularly discloses an image-text delivery sorting management system which scans a two-dimensional code of an image-text commodity to be delivered through a scanning device to obtain commodity information of the commodity, and collects an image-text appearance image of the image-text commodity to be delivered; and an artificial intelligence algorithm based on deep learning is further introduced to perform semantic analysis on the image-text appearance image and the commodity information of the commodity, so that the commodity information of the commodity is verified based on the appearance characteristics of the image-text commodity to be delivered, and after the commodity information is confirmed to be correct, the commodity information is sent to a delivery sorting center; and automatically transmitting and sorting the to-be-delivered image-text commodities according to the commodity information. Through the verification mechanism, the consistency of commodity input information and real objects can be effectively ensured, the error rate is reduced, and accurate identification and efficient sorting of image-text commodities are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image recognition technology, and more specifically, to a graphic and text outbound sorting management system. Background Art

[0002] In modern logistics and warehouse management systems, outbound sorting management is a critical step in ensuring the accurate and efficient delivery of goods from the warehouse to the customer. With the rapid development of the e-commerce industry, consumers' expectations for delivery speed and service quality continue to rise. This poses significant challenges to warehousing and logistics service providers, especially when handling a large number of different types of graphic and text products.

[0003] Traditional outbound sorting methods rely primarily on manual operations or automated systems based on barcodes or QR codes. Manual sorting is not only inefficient but also prone to human error. While existing automated sorting systems improve sorting speed, their accuracy relies on upfront data entry. Any discrepancies between data and physical items, such as incorrect product information entry or product mix-ups, are difficult to detect and correct in a timely manner, leading to customers receiving products that do not match their orders.

[0004] In addition, in the field of graphic and text products, as the market demand for personalized customized products grows, more and more products have unique graphic and text appearance designs. Although this meets the diverse needs of consumers, it undoubtedly adds additional complexity to the early data entry process and the subsequent sorting operations.

[0005] Therefore, in order to ensure the accuracy of the sorting of graphic and text products outbound, an optimized graphic and text outbound sorting management system is expected. Summary of the Invention

[0006] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a graphic and text outbound sorting management system, which scans the QR code of the graphic and text goods to be shipped by a scanning device to obtain its product information, and at the same time collects the graphic and text appearance image of the graphic and text goods to be shipped, and further introduces an artificial intelligence algorithm based on deep learning to perform semantic analysis on the graphic and text appearance image and product information of the goods, so as to verify its product information based on the appearance characteristics of the graphic and text goods to be shipped, and then, after confirming that the product information is correct, the product information is sent to the outbound sorting center, and the graphic and text goods to be shipped are automatically transmitted and sorted according to the product information. Through this verification mechanism, the consistency between the product input information and the actual object can be effectively ensured, the error rate can be reduced, and the accurate identification and efficient sorting of graphic and text goods can be achieved.

[0007] Accordingly, according to one aspect of the present application, a graphic and text outbound sorting management system is provided, which includes:

[0008] The product information acquisition module is used to scan the QR code of the product to be shipped through a scanning device to obtain product information;

[0009] A graphic appearance image acquisition module, configured to acquire the graphic appearance image of the graphic product to be shipped by using the scanning device;

[0010] A commodity information verification module, configured to verify the commodity information based on the image of the commodity to be shipped to obtain a verification result;

[0011] A commodity information sending module, configured to send the commodity information to a delivery sorting center in response to the verification result indicating that the commodity information is correct;

[0012] A sorting management module is configured to determine, at the outbound sorting center, a target sorting exit location based on the commodity information and generate a sorting instruction; the outbound sorting center transmits the sorting instruction to an automated sorting device so that the automated sorting device places the to-be-outbound graphic commodity on a correct conveyor belt, and the conveyor belt transports the to-be-outbound graphic commodity to the target sorting exit;

[0013] The commodity information verification module includes:

[0014] A graphic image simulation generation unit, configured to generate a graphic image based on the product information to obtain a graphic product generated image;

[0015] An image feature extraction unit is used to extract image features from the generated image of the graphic product and the graphic appearance image of the graphic product to be shipped to obtain a semantic coding vector of the generated image of the graphic product and a semantic coding vector of the graphic appearance image;

[0016] A semantic alignment interaction unit, configured to perform feature alignment interaction based on semantic information guidance on the image semantic coding vector generated for the graphic product and the image semantic coding vector for the graphic appearance to obtain a simulated-real semantic alignment coding vector for the graphic product;

[0017] The verification result generating unit is used to determine the verification result based on the image-text product simulation-real semantic alignment coding vector.

[0018] Preferably, the graphic image simulation generation unit includes:

[0019] a product information semantic encoding subunit, configured to perform semantic encoding on the product information to obtain a product information semantic encoding vector;

[0020] The product image generation subunit is used to input the product information semantic encoding vector into a graphic image generator based on a generative adversarial network to obtain the graphic product generated image.

[0021] Preferably, the commodity information semantic coding subunit is used to:

[0022] The product information is semantically encoded using a semantic encoder based on the Bert model to obtain a semantic encoding vector of the product information.

[0023] Preferably, the image feature extraction unit is used to:

[0024] The generated image of the graphic product and the graphic appearance image of the graphic product to be shipped are input into a graphic product twin detection network comprising a first convolutional neural network model and a second convolutional neural network model to obtain a semantic coding vector of the generated image of the graphic product and a semantic coding vector of the graphic appearance image.

[0025] Preferably, the semantic alignment interaction unit includes:

[0026] A semantic offset analysis subunit, configured to perform semantic offset analysis on the image semantic coding vector generated for the graphic product and the image semantic coding vector for the graphic appearance to obtain a simulated-real fine-grained semantic information field for the graphic product;

[0027] The semantic alignment coding subunit is used to perform feature mapping semantic alignment interaction on the image semantic coding vector generated by the image text product and the image semantic coding vector of the image text appearance based on the simulated-real fine-grained semantic information field of the image text product to obtain the simulated-real semantic alignment coding vector of the image text product.

[0028] Preferably, the semantic shift analysis subunit is used to:

[0029] Inputting the semantic coding vector of the generated image of the graphic and text product and the semantic coding vector of the graphic and text appearance image into a dimensional modulation module based on point convolution to obtain a modulated semantic coding vector of the generated image of the graphic and text product and a modulated semantic coding vector of the graphic and text appearance image, wherein the modulated semantic coding vector of the generated image of the graphic and text product and the modulated semantic coding vector of the graphic and text appearance image have the same feature dimension;

[0030] After fine-grained association coding is performed on the modulated graphic product generation image semantic coding vector and the modulated graphic appearance image semantic coding vector, they are input into a convolution kernel-based semantic information field coding network to obtain the graphic product simulation-real fine-grained semantic information field.

[0031] Preferably, the semantic alignment encoding subunit is used to:

[0032] Mapping the modulated image-text product generated image semantic coding vector and the modulated image-text appearance image semantic coding vector to the image-text product simulation-real fine-grained semantic information field to obtain a fine-grained aligned image-text product generated image semantic coding vector and a fine-grained aligned image-text appearance image semantic coding vector respectively;

[0033] A position-weighted sum of the semantic coding vector of the fine-grained aligned image-text product generation image and the semantic coding vector of the fine-grained aligned image-text appearance image is calculated to obtain the image-text product simulation-real semantic alignment coding vector.

[0034] Preferably, the verification result generating unit is used to:

[0035] The image-text commodity simulation-real semantic alignment coding vector is input into a commodity information verification module based on a classifier to obtain the verification result.

[0036] This application has at least the following technical effects:

[0037] Compared with the existing technology, the graphic and text outbound sorting management system provided by this application uses a scanning device to scan the QR code of the graphic and text products to be shipped to obtain their product information, and at the same time collects the graphic and text appearance images of the graphic and text products to be shipped, and further introduces an artificial intelligence algorithm based on deep learning to perform semantic analysis on the graphic and text appearance images and product information of the products to be shipped, thereby verifying their product information based on the appearance characteristics of the graphic and text products to be shipped. After confirming that the product information is correct, the product information is sent to the outbound sorting center, and the graphic and text products to be shipped are automatically transmitted and sorted based on the product information. Through this verification mechanism, the consistency between the product input information and the actual product can be effectively ensured, the error rate can be reduced, and the accurate identification and efficient sorting of graphic and text products can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0039] Figure 1 This is a block diagram of a graphic and text outbound sorting management system according to an embodiment of the present application.

[0040] Figure 2 This is a data flow diagram of the image and text outbound sorting management system according to an embodiment of the present application.

[0041] Figure 3This is a block diagram of a commodity information verification module in a graphic and text outbound sorting management system according to an embodiment of the present application.

[0042] Figure 4 This is a block diagram of a graphic image simulation generation unit in a graphic outbound sorting management system according to an embodiment of the present application.

[0043] Figure 5 This is a block diagram of a semantic alignment interaction unit in a picture and text outbound sorting management system according to an embodiment of the present application. DETAILED DESCRIPTION

[0044] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0045] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0046] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0047] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0048] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0049] As mentioned in the above background technology, although the existing automated sorting system has improved the sorting speed, its accuracy of product information depends on the early data entry link. Once the data does not match the actual object, such as product information entry errors, product confusion, etc., it is difficult to detect and correct it in time, which leads to the problem that the products received by customers do not match the orders. In response to the above technical problems, the present application proposes an optimized graphic and text outbound sorting management system, which scans the QR code of the graphic and text products to be shipped by a scanning device to obtain its product information, and at the same time collects the graphic and text appearance images of the graphic and text products to be shipped, and further introduces an artificial intelligence algorithm based on deep learning to perform semantic analysis on the graphic and text appearance images and product information of the products, so as to verify its product information based on the appearance characteristics of the graphic and text products to be shipped, and then, after confirming that the product information is correct, the product information is sent to the outbound sorting center, and the graphic and text products to be shipped are automatically transmitted and sorted according to the product information. Through this verification mechanism, the consistency between the product entry information and the actual object can be effectively ensured, the error rate can be reduced, and the accurate identification and efficient sorting of graphic and text products can be achieved.

[0050] Figure 1 This is a block diagram of a graphic and text outbound sorting management system according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the image and text outbound sorting management system according to the embodiment of the present application. Figure 1 and Figure 2 As shown, the image and text outbound sorting management system 100 includes: a product information acquisition module 110, which is used to scan the QR code of the image and text product to be outbound through a scanning device to obtain product information; a image and text appearance image acquisition module 120, which is used to collect the image and text appearance image of the image and text product to be outbound through the scanning device; a product information verification module 130, which is used to verify the product information based on the image and text appearance image of the image and text product to be outbound to obtain a verification result; a product information sending module 140, which is used to send the product information to the outbound sorting center in response to the verification result that the product information is correct; a sorting management module 150, which is used to determine the target sorting outlet position based on the product information and generate a sorting instruction at the outbound sorting center, and the outbound sorting center transmits the sorting instruction to the automated sorting equipment so that the automated sorting equipment places the image and text product to be outbound on the correct conveyor belt, and the conveyor belt transports the image and text product to be outbound to the target sorting outlet.

[0051] In the aforementioned image and text outbound sorting management system, the product information acquisition module 110 is used to obtain product information by scanning the QR code of the image and text product to be outbound using a scanning device. It should be understood that a QR code can store a large amount of information, including but not limited to the product number, name, size, specifications, batch, etc. By using a scanning device to scan the product's QR code, detailed product information can be quickly and accurately obtained, providing reliable data support for subsequent sorting operations.

[0052] To accurately read QR codes on graphic products waiting to be shipped, the first step is to select the right scanning device. Modern scanning devices come in a wide variety, including fixed-mount scanners, handheld scanners, mobile computers (such as PDAs), and smartphones. For large warehouses or logistics centers, fixed-mount scanners or industrial-grade handheld scanners are preferred because they offer higher reading speeds and stability, making them suitable for handling large volumes of fast-moving goods. Mobile computers, on the other hand, are ideal for work scenarios requiring flexibility, such as when field operators need to carry equipment for inspections or make temporary adjustments to sorting routes. Furthermore, given the potential for dust and humidity in warehouse environments, scanning devices with a high level of protection should be selected to ensure long-term, stable performance. After selecting a scanning device, it is necessary to properly configure and calibrate it to ensure it can recognize QR codes in specific formats and operate stably under varying lighting conditions and ambient conditions. Given the potential for dust and humidity in warehouse environments, scanning devices with a high level of protection should be selected to ensure long-term, stable performance.

[0053] To ensure that every graphic product waiting to be shipped has a unique and well-formatted QR code, first of all, each product should be assigned a unique QR code to avoid confusion caused by duplication. Secondly, fault tolerance is equally important. The QR code should have a certain degree of fault tolerance and can be correctly parsed even if it is partially damaged. The design of information capacity is also crucial. The amount of information contained in the QR code should be determined according to actual needs. It should neither be too much to affect decoding efficiency nor too little to lose tracking significance. In addition, printing quality must be guaranteed to ensure that the QR code is printed with high clarity and clear color contrast to facilitate scanning equipment to accurately capture the image and decode it. In terms of encoding, key information such as product number, specification model, batch number, etc. are usually encoded into the QR code according to predetermined rules. At the same time, auxiliary information such as timestamps and location identifiers can also be added for subsequent data analysis and management optimization.

[0054] Once everything is ready, the scanning process can begin. The product to be shipped is placed in the scanning area, and the scanning device uses its camera or laser to focus on the QR code on the product. Once a valid QR code pattern is detected, the scanner immediately captures the image and quickly decodes it using a built-in algorithm. Upon successful decoding, the relevant product information is converted into a digital signal and transmitted to a backend server or local database for further processing.

[0055] In the above-mentioned graphic and text outbound sorting management system, the graphic and text appearance image acquisition module 120 is used to collect the graphic and text appearance image of the graphic and text goods to be shipped through the scanning device. It should be understood that considering that relying solely on the information provided by the QR code may cause the product information to be inconsistent with the actual product due to data entry errors or other reasons, resulting in sorting errors. Therefore, the present application further collects the graphic and text appearance image of the graphic and text goods to be shipped through the scanning device, compares and verifies the actual appearance characteristics of the goods with the product information contained in the QR code, and further confirms the true status of the goods before sorting, so as to improve the accuracy and reliability of the sorting process.

[0056] First, when choosing a camera, consider both hardware specifications and functional features. In terms of hardware, high resolution (at least 2 megapixels), a moderate frame rate (30fps or higher), and an appropriate focal length and viewing angle are basic requirements. For most applications, these parameters ensure that subtle surface features of the product can be captured while also meeting the needs of rapid, continuous shooting. Considering the potential impact of factors such as dust and humidity in warehouse environments, choosing an industrial-grade camera with a good protection rating (such as IP67) can ensure long-term, stable performance. In terms of functional features, the autofocus function enables the camera to quickly adjust its focus, ensuring clear and sharp images with every shot; cameras with strong low-light performance or built-in fill lights can maintain good image quality in low-light conditions; and cameras that support multispectral imaging technology use light of different wavelengths to illuminate the surface of the product, obtaining more layers of information, such as texture and material, which is particularly important for products that need to capture complex patterns.

[0057] Adjusting the installation position and angle is also crucial. Fixed cameras are generally recommended to be installed directly above the assembly line to cover the entire surface of the product. The angle should be adjusted so that the captured image is as perpendicular to the product plane as possible to avoid image distortion due to tilt. Use a sturdy bracket and mounting hardware to mount the camera to ensure it is not displaced by vibration or other external factors. For work scenarios that require frequent movement or adjustment of the camera position, handheld scanners offer greater operational flexibility. Their lightweight and ergonomic design can reduce operator fatigue after prolonged use. Also, ensure that the wireless or wired connection between the handheld scanner and the backend system is stable and reliable to avoid data transmission interruption or loss.

[0058] Calibration of the lighting system is equally critical. Choose the appropriate light source type, such as LED lights, ring lights, or backlights, based on the reflective properties of the product's surface. For brightly colored or highly reflective objects, ring lights provide even illumination; for transparent or translucent items, backlights can better highlight contours. The color temperature is generally around 5000K to ensure high color reproduction. Additionally, ensure that the light emitted by the light source is evenly distributed within the shooting area to reduce shadows and dark corners, and improve the lighting effect by adjusting the light source position and adding auxiliary lighting. It is also crucial to set the appropriate light intensity based on the characteristics of each product. Excessive light may lead to overexposure, while too little light will blur the image. Therefore, it is necessary to have a dimmable device that can flexibly adapt to various shooting needs.

[0059] When it comes to software configuration, it's crucial to ensure that image processing algorithms can operate efficiently and are capable of quickly and accurately identifying and analyzing images. This involves optimizing features like image resolution, color depth, and focus adjustment to suit the characteristics of different products. For example, products with bright colors and complex patterns may require higher resolution and more sophisticated color processing; whereas, for products with simple packaging, lower parameter requirements can be employed to speed up processing.

[0060] Throughout the scanning process, a well-designed user interface plays an indispensable role. Intuitive operating instructions, real-time status displays, and clear and concise result feedback help improve work efficiency and reduce human error. For example, when the system detects an error in product information, a prompt box will immediately pop up on the interface to guide the operator on how to resolve the issue. The system also records historical data for each scan, facilitating future queries and audits. For issues that cannot be resolved through automatic comparison, the system switches to manual intervention mode. The operator can carefully examine the provided image and related information and decide whether to retake the image or take other remedial measures. Furthermore, a problem tracking mechanism can be established to record the handling process and results of each abnormal case, providing a reference for future improvements.

[0061] When all preparations are completed, the specific acquisition steps can be implemented. First, the graphic products to be shipped out are placed in the scanning area, and the operator needs to ensure that the products are within the predetermined shooting range. For some products with special shapes or larger sizes, auxiliary tools (such as trays, brackets, etc.) may be needed to stabilize their position to prevent movement or tilting during shooting. In addition, dust, stains and other debris on the surface of the products need to be cleaned to ensure that the image quality is not affected. Before the formal shooting, a preprocessing step is usually performed, that is, a low-resolution rapid scan is performed to obtain the approximate outline and location information of the product. This not only helps determine the best shooting angle, but also can detect and correct possible problems in advance, such as improper placement of products, label obstruction, etc. The preprocessing step helps to improve the success rate and efficiency of subsequent high-resolution image acquisition.

[0062] After the pre-processing is completed, the formal graphic appearance image capture stage begins. The camera starts to aim at a specific part of the product, which usually includes but is not limited to the main visual feature area of the product, such as the brand logo, product name, packaging design, etc. Depending on the specific characteristics of the product, it may be necessary to adjust the camera's angle, focal length and other parameters to ensure that the clearest and sharpest image is captured. For products with complex patterns or multi-layered structures, it may be necessary to take photos from different angles multiple times to fully cover all important details. For example, for products with a three-dimensional relief effect, they can be photographed from multiple sides to fully record their three-dimensional features. In addition, multispectral imaging technology can be used to illuminate the surface of the product with light of different wavelengths to obtain more levels of information, such as texture, material, etc.

[0063] Finally, all collected graphic appearance images and related data are properly saved to form a complete image database.

[0064] In the above-mentioned image and text outbound sorting management system, the commodity information verification module 130 is used to verify the commodity information based on the image and text appearance image of the image and text commodity to be outbound to obtain a verification result. Figure 3 FIG is a block diagram of a commodity information verification module in a graphic and text outbound sorting management system according to an embodiment of the present application. Figure 3As shown, the product information verification module 130 includes: a graphic image simulation generation unit 131, which is used to generate a graphic image based on the product information to obtain a graphic product generation image; an image feature extraction unit 132, which is used to extract image features of the graphic product generation image and the graphic appearance image of the graphic product to be shipped out to obtain a graphic product generation image semantic coding vector and a graphic appearance image semantic coding vector; a semantic alignment interaction unit 133, which is used to perform feature alignment interaction based on semantic information guidance on the graphic product generation image semantic coding vector and the graphic appearance image semantic coding vector to obtain a graphic product simulation-real semantic alignment coding vector; a verification result generation unit 134, which is used to determine the verification result based on the graphic product simulation-real semantic alignment coding vector.

[0065] Specifically, the graphic image simulation generation unit 131 is used to generate a graphic image based on the product information to obtain a graphic product generation image. Specifically, during the verification process, considering that the product information stored in the QR code is in text form, and the graphic appearance image is in image form, there is a difference in the information expression form between the two. Therefore, in order to achieve effective comparative verification between the two, the present application further generates a graphic image based on the product information to obtain a graphic product generation image corresponding to the product information, so as to facilitate comparative analysis between images. Among them, Figure 4 FIG is a block diagram of a graphic image simulation generation unit in a graphic and text outbound sorting management system according to an embodiment of the present application. Figure 4 As shown, the graphic image simulation generation unit 131 includes: a product information semantic encoding subunit 1311, which is used to semantically encode the product information to obtain a product information semantic encoding vector; and a product image generation subunit 1312, which is used to input the product information semantic encoding vector into a graphic image generator based on a generative adversarial network to obtain the graphic product generated image.

[0066] Specifically, the product information semantic encoding subunit 1311 is used to semantically encode the product information to obtain a product information semantic encoding vector. It should be understood that in order to convert the product information in text form into a numerical representation that can be understood by a computer, the present application further uses natural language processing (NLP) technology to semantically encode the product information. In a specific example of the present application, a semantic encoder based on the Bert model is used to semantically encode the product information to obtain the product information semantic encoding vector. Those skilled in the art should know that the Bert model is a pre-trained language representation model. Unlike traditional unidirectional language models, the Bert model adopts a bidirectional encoder structure and can simultaneously consider contextual information, thereby generating a richer and more accurate word embedding representation. Through bidirectional encoding and self-attention mechanism, the Bert model can capture complex semantic information in product descriptions. For example, for product descriptions containing multiple attributes (such as color, size, brand, etc.), the Bert model can understand the relationship between each attribute and provide a more accurate semantic representation. Compared with traditional keyword matching-based methods, the semantic encoding vector generated by the Bert model can better reflect the true meaning of product information. This allows accurate identification of similar products even when descriptions differ when comparing product information, reducing false positives. When sorting products with images and text for shipment, basic product information is obtained by scanning QR codes, and then the product descriptions are semantically encoded using the BERT model. This not only verifies the consistency of product information but also detects potential issues such as mismatched descriptions or missing key attributes.

[0067] Specifically, the product image generation subunit 1312 is used to input the product information semantic encoding vector into a graphic image generator based on a generative adversarial network to obtain the graphic product generated image. That is, in order to convert the abstract product description into a concrete, visual image, the present application uses a generative adversarial network to process the product information semantic encoding vector, so as to utilize the powerful image generation capability of the generative adversarial network to realize the mapping from the text semantic space to the image space, and obtain the graphic product generated image. Those skilled in the art should know that a generative adversarial network (GAN) consists of a generator and a discriminator, which are trained together in a mutually adversarial manner to optimize the quality of the generated image. The generator receives the product information semantic encoding vector as input, and attempts to generate realistic graphic product images based on this information; the discriminator is responsible for distinguishing between the generated fake images and the real product images, and outputs a probability value indicating the possibility of true or false. As training progresses, the generator attempts to deceive the discriminator, making it difficult to distinguish between generated images and real images, while the discriminator continues to learn to improve its recognition ability. Both are gradually optimized in this process, and ultimately the generator can produce high-quality and realistic images of graphic products.

[0068] Specifically, the image feature extraction unit 132 is used to: input the image of the graphic product and the graphic appearance image of the graphic product to be shipped into a graphic product twin detection network including a first convolutional neural network model and a second convolutional neural network model to obtain the semantic coding vector of the graphic product generation image and the semantic coding vector of the graphic appearance image. Specifically, the present application further extracts image features from the graphic product generation image and the graphic appearance image of the graphic product to be shipped, so as to realize the verification of the product information based on the comparative analysis between the image features. Here, in order to avoid additional errors introduced due to inconsistent encoding methods, the present application adopts a graphic product twin detection network to process the graphic product generation image and the graphic appearance image of the graphic product to be shipped. Among them, the image-text product twin detection network includes a first convolutional neural network model and a second convolutional neural network model, which are respectively used to process the image-text product generation image and the image-text appearance image, so as to capture key visual information in the image, such as shape, color, texture, etc., through sliding convolution operations, and generate corresponding semantic coding vectors of the image-text product generation image and the image-text appearance image. In addition, the first convolutional neural network model and the second convolutional neural network model have the same network structure and parameter settings. Through this design, it can be ensured that the two network models have the same feature extraction capabilities when processing the image-text product generation image and the image-text appearance image, ensuring the consistency of the encoding process, so that the extracted image features are comparable.

[0069] Specifically, the semantic alignment interaction unit 133 is used to perform feature alignment interaction based on semantic information guidance on the semantic coding vector of the graphic product generation image and the semantic coding vector of the graphic appearance image to obtain a graphic product simulation-real semantic alignment coding vector. It should be understood that, considering that the graphic product generation image and the graphic appearance image come from different domains (i.e., different information sources), the former is an image generated based on text information, and the latter is an actual photographed product appearance image. There may be significant data distribution differences between the two, which makes direct position-by-position feature interaction produce information differences. In this regard, the present application proposes a feature alignment interaction method based on semantic information guidance, which constructs a semantic information field by learning the feature differences between the semantic coding vector of the graphic product generation image and the semantic coding vector of the graphic appearance image, so that image features from different domains can be effectively aligned and interacted in the shared semantic information field, thereby significantly reducing the feature extraction error caused by image domain differences and improving the matching accuracy between the graphic appearance image and the graphic product generation image. Among them, Figure 5 FIG is a block diagram of a semantic alignment interaction unit in a picture and text outbound sorting management system according to an embodiment of the present application. Figure 5 As shown, the semantic alignment interaction unit 133 includes: a semantic offset analysis subunit 1331, which is used to perform semantic offset analysis on the image semantic coding vector generated by the graphic product and the image semantic coding vector of the graphic appearance to obtain a graphic product simulation-real fine-grained semantic information field; a semantic alignment coding subunit 1332, which is used to perform feature mapping semantic alignment interaction on the image semantic coding vector generated by the graphic product and the image semantic coding vector of the graphic appearance based on the graphic product simulation-real fine-grained semantic information field to obtain the graphic product simulation-real semantic alignment coding vector.

[0070] More specifically, the semantic shift analysis subunit 1331 is configured to: first, input the semantic coding vector of the generated image of the graphic and text product and the semantic coding vector of the graphic and text appearance image into a dimension modulation module based on point convolution to obtain a modulated semantic coding vector of the generated image of the graphic and text product and a modulated semantic coding vector of the graphic and text appearance image, wherein the modulated semantic coding vector of the generated image of the graphic and text product and the modulated semantic coding vector of the graphic and text appearance image have the same feature dimension, which can be expressed as follows:

[0071] v′1=Leaky ReLU{Conv 1×1 (v1)}

[0072] v′2=Leaky ReLU{Conv 1×1 (v2)}

[0073] Among them, v1 represents the semantic coding vector of the image generated by the graphic product, v2 represents the semantic coding vector of the graphic appearance image, Conv 1×1 (·) represents the point convolution operation,

[0074] Leaky ReLU is a leaky linear rectifier function, v′1 represents the semantic coding vector of the image generated by the image and text after modulation, and v′2 represents the semantic coding vector of the image appearance after modulation.

[0075] That is, in order to ensure that the semantic coding vector of the image generated by the graphic product and the semantic coding vector of the graphic appearance image have the same feature dimension, the two are first dimensionally modulated by performing point convolution processing on the two to achieve preliminary dimensional alignment.

[0076] Then, after fine-grained associative coding is performed on the semantic coding vector of the modulated graphic product image and the semantic coding vector of the modulated graphic appearance image, the vector is input into the semantic information field coding network based on the convolution kernel to obtain the graphic product simulation-real fine-grained semantic information field. Specifically, after the feature dimension is adjusted, the present application further captures the subtle semantic relationship between the two by performing fine-grained associative coding on the semantic coding vector of the modulated graphic product image and the semantic coding vector of the modulated graphic appearance image. Then, by stacking multiple convolutional layers, feature extraction is performed based on the local receptive field of the convolution kernel to mine the position alignment and offset information between the semantic coding vector of the graphic product image generation and the semantic coding vector of the graphic appearance image, and gradually learn the mapping from low-level features to high-level semantic concepts, thereby constructing a graphic product simulation-real fine-grained semantic information field, providing precise guidance for subsequent feature mapping.

[0077] The above process is expressed as follows:

[0078]

[0079] in,(·) T represents the transpose of a vector, represents a matrix multiplication operation, L is the characteristic scale value of the semantic coding vector of the image generated by the modulated graphic product and the semantic coding vector of the modulated graphic appearance image, M x Represents the image and text product simulation-real fine-grained semantic association encoding matrix, Conv 3×3 (·) represents a 3×3 convolution operation, and Ω represents the simulated-real fine-grained semantic information field of the image and text product.

[0080] More specifically, the semantic alignment coding subunit 1332 is configured to: first, map the modulated image-text product generated image semantic coding vector and the modulated image-text appearance image semantic coding vector to the image-text product simulated-real fine-grained semantic information field to obtain a fine-grained aligned image-text product generated image semantic coding vector and a fine-grained aligned image-text appearance image semantic coding vector, respectively, which can be expressed as follows:

[0081]

[0082] Among them, v 1t and v 2t They represent the image semantic encoding vector generated by fine-grained alignment of image and text products and the image semantic encoding vector generated by fine-grained alignment of image and text appearance.

[0083] That is, the semantic coding vector of the image generated by the graphic and text product and the semantic coding vector of the graphic and text appearance image after the above-mentioned dimension modulation are subjected to cross-domain feature conversion, so as to respectively map the two to the graphic and text product simulation-real fine-grained semantic information field, and realize semantic alignment between the two. During the mapping process, the various position features in the modulated semantic coding vector of the graphic and text product generated image and the modulated semantic coding vector of the graphic and text appearance image will be adjusted according to the semantic strength and correlation of the corresponding position in the graphic and text product simulation-real fine-grained semantic information field, so as to better match the semantic distribution in the semantic information field, thereby obtaining the fine-grained aligned semantic coding vector of the graphic and text product generated image and the fine-grained aligned semantic coding vector of the graphic and text appearance image. In this way, not only the original semantic information of the graphic and text product generated image and the graphic and text appearance image is retained, but also the two have the fine-grained position alignment characteristic of semantic information, that is, the semantic consistency between each other is enhanced.

[0084] Then, the position-weighted sum of the semantic coding vector of the fine-grained aligned image-text product generation image and the semantic coding vector of the fine-grained aligned image-text appearance image is calculated to obtain the image-text product simulation-real semantic alignment coding vector, which is expressed as:

[0085] v c =αv 1t +βv 2t

[0086] Among them, α and β are learnable weight parameters, v c Represents the simulated-real semantic alignment encoding vector of the image and text product.

[0087] Specifically, after semantic alignment of features, a position-weighted sum calculation is performed to achieve semantic interaction and fusion between the generated product image and the image-text appearance image, providing precise information guidance for subsequent verification steps. This ensures a close semantic connection between the generated product image and the image-text appearance image, enabling accurate identification and verification of differences between the two.

[0088] Specifically, the verification result generating unit 134 is used to input the simulated-real semantic alignment coding vector of the image-text product into a product information verification module based on a classifier to obtain the verification result. In a specific example of the present application, the classifier is trained by supervised learning to accurately identify the image difference features contained in the simulated-real semantic alignment coding vector of the image-text product, and converts it into a judgment of the correctness or error of the product information through a Softmax function to obtain the final verification result, thereby guiding the product's outbound decision.

[0089] Taking into account that the semantic coding vector of the image generated by the graphic product and the semantic coding vector of the graphic appearance image respectively represent the image semantic coding features of the graphic product generated image determined by adversarial generation of product information and the image semantic coding features of the graphic appearance image of the real graphic product to be shipped, during the alignment interaction between features based on the semantic information field, the simulation of the graphic product generated image will cause a relatively significant non-content dimension difference between it and the graphic appearance image of the graphic product to be shipped, such as image format difference, resolution difference, etc., which makes the graphic product simulation-real semantic alignment coding vector have probabilistic convergence and divergence based on different feature interactions, thereby affecting the accuracy of the verification result obtained by the product information verification module based on the classifier.

[0090] Preferably, inputting the simulated-real semantic alignment coding vector of the image-text product into a classifier-based product information verification module to obtain a verification result includes:

[0091] Inputting the simulated-real semantic alignment coding vector of the product image and text into the product information verification module based on the classifier to determine a verification likelihood estimate corresponding to the simulated-real semantic alignment coding vector of the product image and text, wherein the verification likelihood estimate represents a probability value that the verification result is false;

[0092] Calculate the calibration likelihood compensation component of the calibration likelihood estimate to obtain a calibration inverse likelihood estimate, which is expressed as:

[0093] α=1-p

[0094] Where p represents the check likelihood estimate, and α represents the check inverse likelihood estimate;

[0095] Calculate the product of the main feature parameters of the image-text product simulation-real semantic alignment coding vector and the verification likelihood estimate value to obtain the likelihood product of the main feature parameters of the first image-text product simulation-real semantic alignment coding, expressed as: i ×p, where v i The main feature parameter representing the simulated-real semantic alignment encoding vector of the image-text product, i.e., the feature value at the i-th position;

[0096] Calculate the product of the auxiliary feature parameters of the image-text product simulation-real semantic alignment coding vector and the verification inverse likelihood estimate value to obtain the inverse likelihood product of the auxiliary feature parameters of the second image-text product simulation-real semantic alignment coding, expressed as: j ×α, where v j The auxiliary feature parameter representing the simulated-real semantic alignment encoding vector of the image-text product, i.e., the feature value at the j-th position;

[0097] Based on the quotient operation of the main feature parameter and the verification inverse likelihood estimate, the inverse likelihood quotient of the main feature parameter of the first image-text product simulation-real semantic alignment encoding is generated, which is expressed as:

[0098]

[0099] Based on the quotient operation of the auxiliary feature parameter and the verification likelihood estimate, a likelihood quotient of the auxiliary feature parameter of the second image-text product simulation-real semantic alignment encoding is generated, which is expressed as:

[0100]

[0101] Combining the likelihood product of the main feature parameters of the first image-text product simulation-real semantic alignment coding, the inverse likelihood product of the auxiliary feature parameters of the second image-text product simulation-real semantic alignment coding, the inverse likelihood quotient of the main feature parameters of the first image-text product simulation-real semantic alignment coding, and the likelihood quotient of the auxiliary feature parameters of the second image-text product simulation-real semantic alignment coding, the image-text product simulation-real semantic alignment coding correction parameter matrix is obtained, which is expressed as:

[0102]

[0103] v j(j≠i) ∈V

[0104] Wherein, V represents the image-text product simulation-real semantic alignment encoding vector, m i,j Represents the value of the (i, j) position in the image-text product simulation-real semantic alignment encoding correction parameter matrix;

[0105] Based on the image-text product simulation-real semantic alignment coding correction parameter matrix, the image-text product simulation-real semantic alignment coding vector is subjected to domain mapping modulation to obtain a corrected image-text product simulation-real semantic alignment coding vector, which is expressed as:

[0106]

[0107] in, represents matrix multiplication, V′ represents the corrected image-text product simulation-real semantic alignment encoding vector;

[0108] The corrected image-text commodity simulation-real semantic alignment encoding vector is input into the classifier-based emotion recognizer to output the emotion analysis result.

[0109] Therefore, in view of the domain-boundary integral correlation of the multi-dimensional feature topological structure of the graphic product simulation-real semantic alignment coding vector, by defining the overall statistical distribution form of the feature set of the graphic product simulation-real semantic alignment coding vector as the topological constraint boundary, the high-dimensional spatial structure of the feature set of the graphic product simulation-real semantic alignment coding vector is fitted with a connected domain representation composed of feature parameter pairs, thereby avoiding the local topological mapping ambiguity generated when the heterogeneous features of the graphic product simulation-real semantic alignment coding vector are mapped to the high-dimensional feature class convergence space, strengthening the iterative synchronization of each local feature distribution in the projection task, and achieving the common agility of the convergence of the probability convergent and divergent feature set of the graphic product simulation-real semantic alignment coding vector to the probability density distribution space. In this way, the accuracy of the verification result obtained by the classifier-based product information verification module of the graphic product simulation-real semantic alignment coding vector is improved.

[0110] In the aforementioned image and text outbound sorting management system, the product information sending module 140 is configured to send the product information to the outbound sorting center in response to the verification result confirming that the product information is correct. Accordingly, if the verification result indicates that the product information is incorrect, an exception handling process is triggered, including but not limited to rescanning the QR code, manually verifying the product information, and updating the product information database to ensure the accuracy of the product information. If the verification result confirms that the product information is correct, the product information is sent to the outbound sorting center for the next outbound operation.

[0111] After verifying the product information, it's crucial to ensure secure and reliable data transmission to the shipping sorting center. To this end, the adoption of advanced data transmission protocols and technologies is crucial. Common protocols include HTTPS (Hypertext Transfer Protocol Secure) and SFTP (Secure File Transfer Protocol), both of which provide encryption mechanisms to prevent data theft or tampering during network transmission. Furthermore, the introduction of SSL / TLS encryption technology further enhances data transmission security. To ensure transmission reliability, redundant transmission paths, such as dual network card backup or a distributed file system, are deployed to ensure that data can reach its destination even if a path fails. In addition to protocol-level safeguards, network bandwidth and latency must also be considered. Logistics warehouses typically process large amounts of data, ensuring that the network infrastructure has sufficient bandwidth to support efficient data transmission. Furthermore, minimizing network latency is crucial, especially in scenarios with high real-time requirements, such as rush order processing. Cloud storage services are also a good option for large-scale data transmission, offering robust scalability and simplifying data management and maintenance.

[0112] To improve work efficiency, manual intervention should be minimized and automated processing should be implemented at every stage, from the completion of product information verification to the final delivery to the outbound sorting center. Specifically, once the product information is confirmed to be correct, the system automatically triggers the delivery operation without the need for human intervention. This not only saves time but also prevents human error. For large quantities of product information, the system should have efficient batch processing capabilities and support sending information on multiple products at once. This helps speed up overall processing, especially during peak periods or large-scale promotional events. Throughout the transmission process, the system should track the status of each product information in real time and provide immediate feedback to the operator. For example, when a product information is successfully sent to the sorting center, a corresponding prompt will be displayed on the interface; if a problem is encountered, a warning box will pop up and a log will be recorded to facilitate subsequent investigation.

[0113] In the above-mentioned graphic and text outbound sorting management system, the sorting management module 150 is used to determine the target sorting outlet location based on the product information and generate sorting instructions at the outbound sorting center. The outbound sorting center transmits the sorting instructions to the automated sorting equipment so that the graphic and text products to be outbound are placed on the correct conveyor belt through the automated sorting equipment, and the conveyor belt transports the graphic and text products to be outbound to the target sorting outlet. Specifically, after receiving accurate product information, the outbound sorting center will automatically select the appropriate conveyor belt and sorting path based on the sorting outlet location information contained in the product information to ensure that the products can be sorted to the designated outlet location accurately. Sorting instructions include but are not limited to key parameters such as product number, target sorting outlet number, and conveyor belt path. In order to ensure the accuracy and executability of the instructions, standardized data formats such as JSON or XML are usually adopted to ensure seamless connection between different systems. In addition, considering the complexity in actual operation, redundant design will also be introduced to reserve processing space for possible abnormal situations.

[0114] The generated sorting instructions are then transmitted to the automated sorting equipment, which is responsible for physically placing the graphic products to be shipped on the correct conveyor belt. This type of equipment is usually composed of multiple modules, including but not limited to robotic arms, automatic guided vehicles (AGVs), conveyor belt systems, etc. Each module is equipped with sensors and control systems that can accurately perform corresponding actions based on the received sorting instructions. For example, the robotic arm can grab the goods and place them on the designated conveyor belt; the AGV can move flexibly within the warehouse to transport the goods to the designated location; the conveyor belt system is responsible for transporting the goods along the predetermined path to the target sorting exit.

[0115] When the automated sorting equipment places the graphic products to be shipped onto the correct conveyor belt according to the sorting instructions, the conveyor belt transports the products to the target sorting exit. The conveyor belt itself needs to have sufficient load-bearing capacity and stability to ensure that the products will not be damaged or shifted during the entire transportation process. At the same time, the speed control of the conveyor belt must also be considered to match the different types of sorting needs. For example, during peak periods or large-scale promotional events, the conveyor belt speed may need to be increased to improve processing efficiency; under normal circumstances, the speed can be appropriately reduced to reduce energy consumption and wear. To further optimize the efficiency of the conveyor belt, it is also possible to introduce an intelligent scheduling algorithm to dynamically adjust the operating status of each conveyor belt section to avoid congestion.

[0116] In summary, the image and text outbound sorting management system based on the embodiment of the present application is explained, which scans the QR code of the image and text product to be shipped by a scanning device to obtain its product information, and at the same time collects the image and text appearance image of the image and text product to be shipped, and further introduces an artificial intelligence algorithm based on deep learning to perform semantic analysis on the image and text appearance image and product information of the product, so as to verify its product information based on the appearance characteristics of the image and text product to be shipped, and then, after confirming that the product information is correct, the product information is sent to the outbound sorting center, and the image and text products to be shipped are automatically transmitted and sorted according to the product information. Through this verification mechanism, the consistency between the product input information and the actual object can be effectively ensured, the error rate can be reduced, and the accurate identification and efficient sorting of image and text products can be achieved.

[0117] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details of the above embodiments are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.

[0118] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0119] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

Claims

1. The image and text outbound sorting management system is characterized by: include: The product information acquisition module is used to scan the QR code of the product to be shipped through a scanning device to obtain product information; A graphic appearance image acquisition module, configured to acquire the graphic appearance image of the graphic product to be shipped by using the scanning device; A commodity information verification module, configured to verify the commodity information based on the image of the commodity to be shipped to obtain a verification result; A commodity information sending module, configured to send the commodity information to a delivery sorting center in response to the verification result indicating that the commodity information is correct; A sorting management module is configured to determine, at the outbound sorting center, a target sorting exit location based on the commodity information and generate a sorting instruction; the outbound sorting center transmits the sorting instruction to an automated sorting device so that the automated sorting device places the to-be-outbound graphic commodity on a correct conveyor belt, and the conveyor belt transports the to-be-outbound graphic commodity to the target sorting exit; The commodity information verification module includes: A graphic image simulation generation unit, configured to generate a graphic image based on the product information to obtain a graphic product generated image; An image feature extraction unit is used to extract image features from the generated image of the graphic product and the graphic appearance image of the graphic product to be shipped to obtain a semantic coding vector of the generated image of the graphic product and a semantic coding vector of the graphic appearance image; A semantic alignment interaction unit, configured to perform feature alignment interaction based on semantic information guidance on the image semantic coding vector generated for the graphic product and the image semantic coding vector for the graphic appearance to obtain a simulated-real semantic alignment coding vector for the graphic product; The verification result generating unit is used to determine the verification result based on the image-text product simulation-real semantic alignment coding vector.

2. The image and text outbound sorting management system according to claim 1 is characterized in that: The graphic image simulation generation unit includes: a product information semantic encoding subunit, configured to perform semantic encoding on the product information to obtain a product information semantic encoding vector; The product image generation subunit is used to input the product information semantic encoding vector into a graphic image generator based on a generative adversarial network to obtain the graphic product generated image.

3. The image and text outbound sorting management system according to claim 2 is characterized in that: The commodity information semantic encoding subunit is used to: The product information is semantically encoded using a semantic encoder based on the Bert model to obtain a semantic encoding vector of the product information.

4. The image and text outbound sorting management system according to claim 3 is characterized in that: The image feature extraction unit is used to: The generated image of the graphic product and the graphic appearance image of the graphic product to be shipped are input into a graphic product twin detection network comprising a first convolutional neural network model and a second convolutional neural network model to obtain a semantic coding vector of the generated image of the graphic product and a semantic coding vector of the graphic appearance image.

5. The image and text outbound sorting management system according to claim 4 is characterized in that: The semantic alignment interaction unit includes: A semantic offset analysis subunit, configured to perform semantic offset analysis on the image semantic coding vector generated for the graphic product and the image semantic coding vector for the graphic appearance to obtain a simulated-real fine-grained semantic information field for the graphic product; The semantic alignment coding subunit is used to perform feature mapping semantic alignment interaction on the image semantic coding vector generated by the image text product and the image semantic coding vector of the image text appearance based on the simulated-real fine-grained semantic information field of the image text product to obtain the simulated-real semantic alignment coding vector of the image text product.

6. The image and text outbound sorting management system according to claim 5 is characterized in that: The semantic shift analysis subunit is used to: Inputting the semantic coding vector of the generated image of the graphic and text product and the semantic coding vector of the graphic and text appearance image into a dimensional modulation module based on point convolution to obtain a modulated semantic coding vector of the generated image of the graphic and text product and a modulated semantic coding vector of the graphic and text appearance image, wherein the modulated semantic coding vector of the generated image of the graphic and text product and the modulated semantic coding vector of the graphic and text appearance image have the same feature dimension; After fine-grained association coding is performed on the modulated graphic product generation image semantic coding vector and the modulated graphic appearance image semantic coding vector, they are input into a convolution kernel-based semantic information field coding network to obtain the graphic product simulation-real fine-grained semantic information field.

7. The image and text outbound sorting management system according to claim 6 is characterized in that: The semantic alignment encoding subunit is used to: Mapping the modulated image-text product generated image semantic coding vector and the modulated image-text appearance image semantic coding vector to the image-text product simulation-real fine-grained semantic information field to obtain a fine-grained aligned image-text product generated image semantic coding vector and a fine-grained aligned image-text appearance image semantic coding vector respectively; A position-weighted sum of the semantic coding vector of the fine-grained aligned image-text product generation image and the semantic coding vector of the fine-grained aligned image-text appearance image is calculated to obtain the image-text product simulation-real semantic alignment coding vector.

8. The image and text outbound sorting management system according to claim 7 is characterized in that: The verification result generating unit is configured to: The image-text commodity simulation-real semantic alignment coding vector is input into a commodity information verification module based on a classifier to obtain the verification result.