Receiving and sending standard acquisition method and apparatus, computer device, and storage medium
By segmenting, screening and extracting features from cross-border logistics texts and using a large language model to generate collection and delivery standards, the difficult problem of refining collection and delivery standards in cross-border logistics is solved, and logistics efficiency and accuracy are improved.
Patent Information
- Application Number
- PCT/CN2025/078094
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-20
- Filing Date
- 2025-02-19
- Publication Date
- 2025-09-25
AI Technical Summary
In cross-border logistics, it is difficult to efficiently extract the specific scope of prohibited items and the relevant materials that companies or individuals need to prepare during the delivery process from the huge and complex customs laws and regulations. Existing technology makes it difficult to achieve this efficiently.
By segmenting, filtering, extracting features, and summarizing the target text, and processing it using a large language model, we can determine the acceptance and delivery standards for items. The specific steps include text segmentation, semantic segmentation, feature attribute extraction, and text concatenation, ultimately generating the acceptance and delivery standards.
It has achieved the efficient extraction of the collection and delivery standards of items from complex laws and regulations, improved the efficiency of cross-border logistics, and ensured the accuracy and completeness of information.
Smart Images

Figure CN2025078094_25092025_PF_FP_ABST
Abstract
Description
Method, device, computer equipment and storage medium for obtaining collection and delivery standards
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 202410338775.3, filed on March 20, 2024, entitled “Method, device, computer equipment and storage medium for obtaining collection and delivery standards,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of text processing technology, and in particular to a method, apparatus, computer equipment and storage medium for obtaining collection and delivery standards. Background Art
[0004] Cross-border logistics constantly face the question of whether an item can be sent from region A to region B. This requires a deep understanding of documents such as various customs laws and regulations. Therefore, how to extract the specific scope of prohibited items from this vast and complex document and summarize the relevant materials that businesses or individuals need to prepare during the delivery process is a pressing issue. Summary of the Invention
[0005] According to various embodiments of the present application, a method, apparatus, computer device, and storage medium for obtaining collection and delivery standards are provided.
[0006] In a first aspect, the present application provides a method for obtaining acceptance and delivery standards. The method comprises: segmenting a target text to obtain multiple text blocks; selecting a target text block from the multiple text blocks; wherein the target text block is a text block related to the acceptance and delivery standards; extracting features from the target text blocks to obtain characteristic attributes of an item; concatenating the target text blocks with the same characteristic attributes to obtain characteristic text blocks; and summarizing the characteristic text blocks to obtain the acceptance and delivery standards for the item.
[0007] In one embodiment, the step of segmenting the target text to obtain multiple text blocks includes: segmenting the target text based on a preset number of words and / or preset symbols to obtain multiple segmented texts; and recursively dividing the multiple segmented text blocks based on semantics to obtain multiple text blocks.
[0008] In one embodiment, the plurality of segmented texts include: a first segmented text and a second segmented text, wherein the second segmented text is adjacent to the first segmented text;
[0009] The recursive segmentation of the plurality of segmented text blocks based on semantics to obtain the plurality of text blocks includes:
[0010] Performing text summarization on the first segmented text to obtain a first summary text;
[0011] Merging the first segmented text, the second segmented text, and the first summary text to obtain a first merged text;
[0012] According to the first summary text, semantic segmentation is performed on the combination of the first segmented text and the second segmented text in the first merged text to obtain a first segmented text and a second segmented text; a first segmented summary text is determined; wherein the first segmented summary text is obtained by performing textual summary on the first segmented text;
[0013] The first segmented text and the first segmented summary text are taken as one text block.
[0014] In one embodiment, after semantic segmentation is performed on the combination of the first segmented text and the second segmented text in the first merged text according to the first summary text to obtain a first segmented text and a second segmented text, the method further includes:
[0015] Determining a second segmentation summary text; wherein the second segmentation summary text is obtained by summarizing the second segmentation text and its corresponding context information;
[0016] Using the second segmented summary text as a new first summary text;
[0017] The original second segmented text and the third segmented text are used as the new first segmented text and the new second segmented text respectively; the third segmented text is adjacent to the original second segmented text;
[0018] Return to the step of merging the first segmented text, the second segmented text, and the first summary text to obtain a first merged text, thereby obtaining a plurality of the text blocks.
[0019] In one embodiment, the step of summarizing the first segmented text to obtain a first summary text includes:
[0020] Inputting summary prompt words into the large language model; wherein the summary prompt words include the first segmented text and summary requirement information; the summary requirement information is used to instruct the large language model to summarize the first segmented text;
[0021] The first summary text output by the large language model according to the summary prompt words is obtained.
[0022] In one embodiment, the step of performing semantic segmentation on the combination of the first segmented text and the second segmented text in the first merged text based on the first summary text to obtain a first segmented text and a second segmented text includes:
[0023] Inputting semantic segmentation prompt words into the large language model; wherein the semantic segmentation prompt words include the first merged text and semantic segmentation information; the semantic segmentation information is used to indicate that based on understanding the first summary text, the first segmented text, and the second segmented text, the combination of the first segmented text and the second segmented text is segmented, and the two segments of text obtained by segmentation are summarized;
[0024] The first segmented text, the second segmented text, the first segmented summary text, and the second segmented summary text output by the large language model according to the semantic segmentation prompt words are obtained.
[0025] In one embodiment, the step of selecting a target text block from the plurality of text blocks includes:
[0026] Inputting a screening prompt word into the large language model; wherein the screening prompt word includes the plurality of text blocks and screening requirement information; the screening requirement information includes output requirement information and output condition information; the output requirement information is used to indicate content information of an output result of the large language model; the output condition information is used to indicate condition information corresponding to whether the large language model outputs a yes or no result;
[0027] The target text block is determined based on the output of the large language model according to the screening prompt words.
[0028] In one embodiment, the output requirement information includes:
[0029] Determine whether the text is related to cross-border collection and delivery, and whether it involves specific requirements for cross-border collection and delivery items or collection and delivery entities, and give a result on whether the text content meets the requirements of the task, indicating the result with yes or no.
[0030] In one embodiment, the output condition information includes:
[0031] The first output condition information is used to indicate that if the text content meets the first preset condition, the determination result is "yes";
[0032] The second output condition information is used to indicate that if the text content meets the second preset condition, the determination result is "no".
[0033] In one embodiment, determining the target text block based on the output of the screening prompt word based on the large language model includes:
[0034] The text block containing "yes" in the output result of the large language model is determined as the target text block.
[0035] In one embodiment, the step of extracting features from the target text block to obtain feature attributes of the item includes:
[0036] Inputting a feature extraction prompt word into the large language model; wherein the feature extraction prompt word includes the target text block and feature requirement information; the feature requirement information includes: extraction requirement information and association judgment information; the extraction requirement information is used to indicate the extraction requirement of the large language model; the association judgment information is used to instruct the large language model to determine whether the current text is associated with a preset item;
[0037] The feature attribute corresponding to each target text block output by the large language model according to the feature extraction prompt word is obtained.
[0038] In one embodiment, the step of summarizing the characteristic text block to obtain the acceptance and delivery standards of the item includes:
[0039] Inputting a text summary prompt word into the large language model; wherein the text summary prompt word includes the characteristic text block and text summary information; the text summary information is used to instruct the large language model to extract the cross-border collection and delivery requirements of a certain item from the text, as well as the constraints that need to be taken into account when extracting the requirements;
[0040] The receiving and sending standards output by the large language model according to the text summary prompt words are obtained.
[0041] In one embodiment, the acceptance and delivery standards include: whether the item is allowed to be exported or imported across borders, and / or information on relevant materials that need to be prepared during the acceptance and delivery process;
[0042] The method further includes: forming a collection and delivery standard library by combining the collection and delivery standards of all articles included in the target text.
[0043] In a second aspect, the present application also provides a device for obtaining acceptance and delivery standards. The device comprises: a text segmentation module for segmenting a target text to obtain multiple text blocks; a text screening module for screening a target text block from the multiple text blocks; wherein the target text block is a text block related to the acceptance and delivery standards; a feature extraction module for extracting features from the target text block to obtain characteristic attributes of the item; a text splicing module for splicing the target text blocks with the same characteristic attributes to obtain a characteristic text block; and a text summarization module for summarizing the characteristic text blocks to obtain the acceptance and delivery standards of the item.
[0044] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method for obtaining the collection and delivery standards described in the embodiment of the first aspect are implemented.
[0045] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method for obtaining the collection and delivery standards described in the embodiment of the first aspect.
[0046] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the conventional technology, the following briefly introduces the drawings required for use in the embodiments or the conventional technology descriptions. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without any creative work.
[0048] FIG1 is a diagram showing an application environment of a method for obtaining collection and delivery standards according to an embodiment;
[0049] FIG2 is a flow chart of a method for obtaining delivery and collection standards in one embodiment;
[0050] FIG3 is a schematic diagram of a process for obtaining a text block in one embodiment;
[0051] FIG4 is a schematic diagram of the steps of obtaining a text block in one embodiment;
[0052] FIG5 is a schematic diagram of a process for obtaining a first summary text in one embodiment;
[0053] FIG6 is a schematic diagram of a process for obtaining a first segmented text and a first segmented summary text in one embodiment;
[0054] FIG7 is a schematic diagram of a process for obtaining a target text block in one embodiment;
[0055] FIG8 is a schematic diagram of a process for obtaining characteristic attributes in one embodiment;
[0056] FIG9 is a schematic diagram of a process for obtaining a collection and delivery standard in one embodiment;
[0057] FIG10 is a module diagram of a device for obtaining receiving and sending standards in one embodiment;
[0058] FIG11 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0059] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0060] The method for obtaining the collection and delivery standards provided in the embodiment of the present application can be applied in the application environment shown in Figure 1. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The terminal 102 can obtain the target text from the server 104 through the communication network, and extract the collection and delivery standards of the items from the target text through the steps of text segmentation, screening, feature extraction, splicing and text summarization. Among them, the terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be implemented as an independent server or a server cluster consisting of multiple servers.
[0061] In one embodiment, as shown in FIG2 , a method for obtaining a collection and delivery standard is provided. The method is described by taking the application of the method to the terminal 102 in FIG1 as an example, and includes the following steps:
[0062] Step S100 , performing text segmentation on the target text to obtain multiple text blocks.
[0063] Specifically, after the terminal 102 obtains the target text from the server 104 via the communication network, it first performs text segmentation on the target text to obtain multiple text blocks. This application is intended to refine the collection and delivery standards of items. Therefore, the target text includes laws and regulations related to cross-border logistics or logistics regulations of other scenarios (such as logistics regulations for fresh food). For ease of understanding, the following is an example of cross-border logistics.
[0064] When segmenting the target text, you can segment by length, keywords, punctuation, semantics, paragraphs, and more, depending on the needs of subsequent text processing. You can also combine multiple segmentation methods to segment the target text. After text segmentation, you can obtain multiple text blocks with different contents.
[0065] Step S200: Filter out a target text block from multiple text blocks.
[0066] Specifically, after obtaining multiple text blocks, the terminal 102 needs to filter out a text block related to the collection and delivery standard from the multiple text blocks and use it as a target text block.
[0067] When screening text blocks, the following methods can be used: keyword screening, that is, screening based on the keywords appearing in the text blocks, and classifying text blocks containing specific keywords; semantic analysis screening, that is, using natural language processing technology to perform semantic analysis on the text, understand the meaning expressed in the text, and then perform screening; text similarity matching, that is, using text similarity algorithms to calculate the similarity of texts, and clustering or filtering texts with higher similarity; topic model, that is, using topic models to extract and classify texts, and cluster or filter texts with similar topics.
[0068] By the above screening method, the target text block can be screened out from the multiple text blocks. It can be understood that the target text block includes multiple text blocks related to the collection and delivery standards.
[0069] Step S300: extract features from the target text block to obtain feature attributes of the object.
[0070] In some embodiments, when extracting features from a target text block, natural language processing techniques may be used to identify entities in the text, thereby obtaining characteristic attributes of the object.
[0071] In some embodiments, each target text block corresponds to a characteristic attribute, and the characteristic attributes of multiple target text blocks can be the same. The characteristic attribute is used to identify the characteristics of the items included in the target text block. For example, the characteristic attribute can be the name of a specific item such as dog, cat, toothpaste, or rice dumpling, or the name of a general category such as food, electronic products, or pharmaceuticals.
[0072] Step S400: splicing target text blocks with the same characteristic attributes to obtain characteristic text blocks.
[0073] Specifically, after obtaining the characteristic attributes corresponding to each target text block, target text blocks with the same characteristic attributes are concatenated to obtain a characteristic text block. A characteristic text block includes all target text blocks related to a certain characteristic attribute.
[0074] Step S500: Summarize the characteristic text blocks to obtain the acceptance and delivery standards of the items.
[0075] Specifically, after obtaining the characteristic text block, the characteristic text block includes all target text blocks related to a certain characteristic attribute. At this time, by summarizing the characteristic text block, the corresponding acceptance and delivery standards of the items can be obtained.
[0076] In some embodiments, each characteristic attribute corresponds to a characteristic text block and, at the same time, to an item's acceptance and delivery standards. Through the above steps, the acceptance and delivery standards for all items included in the target text are obtained, thereby forming a collection of acceptance and delivery standards. During cross-border logistics, the acceptance and delivery standards for the corresponding items can be determined by searching the collection and delivery standards library, making it easier for users to review and prepare relevant information, thereby improving the efficiency of cross-border logistics.
[0077] The above-mentioned method for obtaining the collection and delivery standards performs text segmentation on the target text to obtain multiple text blocks, and then filters out the target text blocks related to the collection and delivery standards from the multiple text blocks. At this time, the target text blocks include the collection and delivery standards of multiple items. Feature extraction is then performed on the target text blocks to obtain the characteristic attributes of the items corresponding to the target text blocks, and different characteristic attributes represent different items. The target text blocks with the same characteristic attributes are then spliced to obtain characteristic text blocks, and one characteristic text block includes all the collection and delivery standards of an item. Finally, the characteristic text blocks are summarized to obtain the summarized and refined collection and delivery standards of the items. The present application can extract the collection and delivery standards for mailing items from the huge and complex customs laws and regulations. The collection and delivery standards include the relevant materials that enterprises or individuals need to prepare during the mailing process, thereby improving the efficiency of cross-border logistics.
[0078] In one embodiment, as shown in FIG3 , in step S100 , the step of segmenting the target text to obtain multiple text blocks includes:
[0079] Step S110 , segmenting the target text based on a preset number of characters and / or preset symbols to obtain a plurality of segmented texts.
[0080] Specifically, when performing text segmentation on the target text, this embodiment can segment the target text according to a preset number of words and / or according to preset symbols. The preset number of words (such as 500 words) can determine the length of the text, and the preset symbols can be line breaks, periods, etc.
[0081] In some embodiments, the target text can be segmented based on both a preset word count and a preset symbol. When the preset word count is 500 and the preset symbol is a line break, the target text can be segmented into each paragraph. If a paragraph has more than 500 words, it can be divided into multiple paragraphs based on the word count. As shown in FIG4 , after the target text is segmented, the multiple segmented texts include: a first segmented text, a second segmented text, ..., and an Nth segmented text.
[0082] Step S120 , performing text summary on the first segmented text to obtain a first summary text.
[0083] Specifically, after obtaining multiple segmented texts, a first segmented text, that is, the first segmented text, is subjected to text summary, thereby obtaining a first summary text. The text summary is used to summarize the main content of the segmented texts.
[0084] Step S130: Merge the first segmented text, the second segmented text, and the first summary text to obtain a first merged text.
[0085] Specifically, after obtaining the first summary text, the first segmented text, the second segmented text, and the first summary text obtained by segmentation are merged to obtain a first merged text, wherein the second segmented text is adjacent to the first segmented text.
[0086] Step S140 , semantic segmentation is performed on the first merged text to obtain a first segmented text and a first segmented summary text.
[0087] Specifically, after obtaining the first merged text, the main content of the first merged text also comes from the first segmented text and the second segmented text. Therefore, when performing semantic segmentation on the first merged text, it is divided into two parts to obtain the first segmented text and the second segmented text. At the same time, during semantic segmentation, the main contents of the two parts are summarized respectively to obtain the first segmentation summary text and the second segmentation summary text.
[0088] Step S150: The first segmented text and the first segmented summary text are taken as text blocks.
[0089] Specifically, the first segmented text obtained after semantic segmentation and the first segmented summary text are finally merged to obtain a text block. The second segmented summary text obtained after semantic segmentation is then merged with the original second segmented text and the third segmented text to obtain a new first merged text. Steps S140 and S150 are then repeated until all text blocks are obtained.
[0090] This embodiment processes the target text by adopting a semantic recursive slicing method, which pays more attention to the semantic integrity and can ensure the accuracy and completeness of the information.
[0091] In one embodiment, as shown in FIG5 , in step S120 , the step of summarizing the first segmented text to obtain a first summary text includes:
[0092] Step S121, input summary prompt words into the large language model;
[0093] Step S122: obtaining a first summary text output by the large language model according to the summary prompt words.
[0094] Specifically, this embodiment uses a large language model to summarize the first segmented text to obtain a first summary text. The summary prompt word includes the first segmented text and summary requirement information. The first summary text is the text in the first segmented text that meets the summary requirement information.
[0095] For example, the first segmented text is first input into the large language model, and then the summary requirement information is input. For example, the summary requirement information is: "Your task is to output a summary of the above content." The output result of the large language model is the first summary text.
[0096] In one embodiment, as shown in FIG6 , in step S140 , the step of semantically segmenting the first merged text to obtain a first segmented text and a first segmented summary text includes:
[0097] Step S141, inputting semantic segmentation prompt words into the large language model;
[0098] Step S142 : obtaining the first segmentation text and the first segmentation summary text output by the large language model according to the semantic segmentation prompt words.
[0099] Specifically, this embodiment uses a large language model to semantically segment the first merged text, thereby obtaining a first segmented text and a first segmented summary text. The semantic segmentation prompt word includes the first merged text and the semantic segmentation information, and the first segmented text is the text in the first merged text that meets the semantic segmentation information.
[0100] Specifically, the first merged text is first input into the large language model, followed by the semantic segmentation information. In some embodiments, since the first merged text includes the first segmented text, the second segmented text, and the first summary text, the first segmented text and the second segmented text are used as "partial original text," and the first summary text is used as auxiliary information for the "partial original text."
[0101] Semantic segmentation information includes: text specification information, which specifies the text data to be processed by the large language model; text segmentation information, which instructs the large language model to perform text segmentation; and text summarization information, which instructs the large language model to perform text summarization. Specifically, the semantic segmentation information is as follows: "Your task: 1. Understand the auxiliary information and the "partial original text" of the "partial original text." 2. Analyze the structure of the "partial original text" and split it into two parts where appropriate. 3. Summarize the main content of the first part and provide the starting and ending line numbers of the first part. 4. For the second part, summarize its main content based on the context and the content of the first part." In this case, step 1 is the text specification information, step 2 is the text segmentation information, and steps 3 and 4 are the text summarization information. The first part after segmentation is the first segmented text, and the second part is the second segmented text. The main content of the first part summarized is the first segmentation summary text, and the main content of the second part summarized is the second segmentation summary text.
[0102] In one embodiment, as shown in FIG7 , in step S200 , the step of selecting a target text block from multiple text blocks includes:
[0103] Step S210: inputting screening prompt words into the large language model;
[0104] Step S220 , obtaining a target text block output by the large language model according to the screening prompt words.
[0105] Specifically, this embodiment uses a large language model to filter multiple text blocks to obtain a target text block.
[0106] The screening prompt word includes multiple text blocks and screening requirement information, and the target text block is the text block among the multiple text blocks that meets the screening requirement information. In some embodiments, the screening requirement information includes: output requirement information for indicating the output result of the large language model; output condition information for indicating the conditions under which the large language model outputs.
[0107] For example, the output requirement is: "Based on the given title, source text summary, and source text content, consider and judge each example to determine whether it is related to import and export, and whether it involves specific requirements for import and export items or the sender (e.g., individual or company). Finally, provide a result, judging whether the content meets the task requirements, and indicate a yes or no answer." The output condition information is: "Judge as 'yes' if the content meets the following conditions: 1. The text clearly states that a certain item cannot be imported or exported; 2. Specific conditions are set for the import or export of a certain item, and these conditions are clearly listed in the text; 3. The text sets specific requirements for companies or individuals shipping specific items, such as requiring certain licenses or qualifications. Judge as 'no' if the content meets the following conditions: 1. The text is completely unrelated to import and export items, such as regulations on pharmacists issuing prescriptions; 2. The text mentions the conditions for importing and exporting a certain item, but does not specify the specific conditions or requirements; 3. The content is indeed related to import and export, but focuses primarily on departmental regulations, procedures, and penalties, rather than specific requirements for the shipped items or senders." In this setting, the large language model outputs whether each text block is related to imports and exports. For example, a text block containing the following text: "Pharmacists must perform four checks and ten comparisons when dispensing prescriptions" is irrelevant to imports and exports, so the corresponding output is "No." In this case, the text block with a "Yes" output is used as the target text block.
[0108] In one embodiment, as shown in FIG8 , in step S300 , the step of extracting features from the target text block to obtain feature attributes of the item includes:
[0109] Step S310, inputting feature extraction prompt words into the large language model;
[0110] Step S320: Obtain the feature attributes output by the large language model based on the feature extraction prompt word.
[0111] Specifically, this embodiment uses a large language model to extract features from a target text block, thereby obtaining characteristic attributes of the object.
[0112] The feature extraction prompt includes a target text block and feature requirement information. The feature attribute is the text in the target text block that meets the feature requirement information. In some embodiments, the feature requirement information includes: extraction requirement information, which indicates the extraction requirements of the large language model; and association determination information, which indicates whether the large language model determines whether the current text is associated with a specific item.
[0113] For example, the extraction requirements input into the large language model are as follows: "Your task is to extract the names of items related to the collection and delivery standards from the text. The requirements are as follows: 1. Minimization: For items that can be disassembled, try to disassemble them. For example, live animals are prohibited from being mailed (except dogs and cats). Here, dogs and cats can be used as common product name categories, so they need to be split into dogs, cats, and live animals (except dogs and cats); 2. No need to split attributes: Different attributes of the same item do not need to be separated. For example, bird's nest is prohibited from being mailed (except for sterilized ones). Here, sterilization is just an attribute of bird's nest, not a product name itself, so there is no need to split it. Bird's nest itself is already minimized; 3. It should not be an abstract item, such as "other" or "items recorded in XX". It must be specific, real-life items such as toothpaste, rice dumplings, cakes, etc. It can also be a large category such as food, electronic products, and pharmaceuticals. 4. "Any" can be used to represent any item, but it is not appropriate to write "any food" or "any electronic product". The latter can be directly represented by "food" or "electronic product." The input association judgment information is: "Simultaneously determine whether each item is related to food: 1. Food in the customs sense includes normal food, beverages, drinks, tea, tea bags, dried goods, seasonings, and health products; 2. If no judgment can be made, it is assumed to be related to food. If the extracted item is arbitrary, it is also considered to be related to food." Under this setting, the output of the large language model is the feature attribute corresponding to each target text block, that is, the name of the item. For example, the content of a target text block is: "(14) Animal carcasses, animal specimens, and animal-derived waste. (15) Soil and organic cultivation medium." After feature extraction using the large language model, the feature attributes obtained are: "Animal carcasses, no; Animal specimens, no; Animal-derived waste, no; Soil, no; Organic cultivation medium, no." The "no" here indicates that the corresponding feature attribute is not related to food. It is understandable that in some embodiments, the feature requirement information only includes: extraction requirement information, and the feature attributes of the item corresponding to the target text block can be obtained by extracting the requirement information.
[0114] In one embodiment, as shown in FIG9 , in step S500 , the step of summarizing the characteristic text block to obtain the acceptance and delivery standards of the item includes:
[0115] Step S510: inputting text summary prompt words into the large language model;
[0116] Step S520: Obtain the collection and delivery standards output by the large language model based on the text summary prompt words.
[0117] Specifically, this embodiment uses a large language model to summarize the feature text blocks, thereby obtaining the acceptance and delivery standards of the items.
[0118] The text summary prompt includes a feature text block and text summary information. The acceptance criteria are the text in the feature text block that meets the text summary information. For example, the feature text block is first input into the large language model, followed by the text summary information. For example, the text summary information might read: "Your task is to extract the export or import requirements for a certain item from the text. When extracting requirements, please remember the following: 1. Only output the most important requirements—the shorter, more intuitive, and simpler the better. 2. Focus solely on the requirements for the item itself, the sender, or the delivery company. Abstract content such as processes, procedures, responsibilities, declarations, and penalties is not necessary." In this setting, the large language model outputs the acceptance criteria for the item corresponding to each feature text block. For example, the content of the feature text block is: "This section mainly contains the relevant regulations on personal inbound and outbound mailings. It includes the definitions of personal use and reasonable quantity, as well as the value limits for items sent to different regions (RMB 800 for area A and RMB 1,000 for area B). It also explains how to handle items that exceed the prescribed limits." After summarizing the text through the large language model, the obtained standard for the acceptance of items is: "Any items sent by individuals out of the country must meet the following requirements: 1. Reasonable personal use; 2. Value limit requirements: For items sent to area A, the value limit is RMB 800 each time; for items sent to area B, the value limit is RMB 1,000 each time. If it is a single and indivisible item, the value limit may be exceeded."
[0119] In one specific embodiment, the method for obtaining acceptance and delivery standards in this application utilizes a large language model for text segmentation, screening, feature extraction, and text summarization. Because the large language model is trained on a large corpus, it can learn a variety of linguistic knowledge and patterns. Consequently, it possesses strong generalization capabilities, adapting to a variety of languages and contexts, and processing a wide variety of text data. Furthermore, it can process text data more naturally and accurately when performing tasks such as text classification and summary generation.
[0120] In one embodiment, in step S500, after obtaining the acceptance and delivery standards of the items, the contents can be manually modified, and some manually provided general rules can be added, or some acceptance and delivery standards can be deleted, so that the contents of the acceptance and delivery standard library are more accurate.
[0121] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0122] Based on the same inventive concept, embodiments of the present application also provide a device for obtaining receiving and sending standards for implementing the aforementioned method for obtaining receiving and sending standards. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the device for obtaining receiving and sending standards provided below can be found in the aforementioned limitations of the method for obtaining receiving and sending standards, and will not be further elaborated here.
[0123] In one embodiment, as shown in FIG10 , a device for obtaining a collection and delivery standard is provided, comprising: a text segmentation module 610 , a text screening module 620 , a feature extraction module 630 , a text splicing module 640 , and a text summarization module 650 , wherein:
[0124] A text segmentation module 610 is used to segment the target text into multiple text blocks;
[0125] The text screening module 620 is used to screen out a target text block from multiple text blocks; wherein the target text block is a text block related to the collection and delivery standards;
[0126] Feature extraction module 630, used to extract features from the target text block to obtain feature attributes of the object;
[0127] The text splicing module 640 is used to splice target text blocks with the same characteristic attributes to obtain characteristic text blocks;
[0128] The text summarizing module 650 is used to perform text summarization on the characteristic text blocks to obtain the acceptance and delivery standards of the items.
[0129] The above-mentioned receiving and sending standards acquisition device obtains multiple text blocks by performing text segmentation on the target text, and then filters out the target text blocks related to the receiving and sending standards from the multiple text blocks. At this time, the target text blocks include the receiving and sending standards of multiple items. The target text blocks are then feature extracted to obtain the characteristic attributes of the items corresponding to the target text blocks, and different characteristic attributes represent different items. The target text blocks with the same characteristic attributes are then spliced to obtain characteristic text blocks, and one characteristic text block includes all the receiving and sending standards of an item. Finally, the characteristic text blocks are summarized to obtain the summarized and refined receiving and sending standards of the items. The present application can extract the receiving and sending standards for mailing items from the huge and complex customs laws and regulations. The receiving and sending standards include the relevant materials that enterprises or individuals need to prepare during the mailing process, thereby improving the efficiency of cross-border logistics.
[0130] In one embodiment, the text segmentation module 610 is also used to segment the target text based on a preset number of words and / or preset symbols to obtain multiple segmented texts; wherein the multiple segmented texts include: a first segmented text and a second segmented text; performing a text summary on the first segmented text to obtain a first summary text; merging the first segmented text, the second segmented text and the first summary text to obtain a first merged text; wherein the second segmented text is adjacent to the first segmented text; performing semantic segmentation on the first merged text to obtain a first segmented text and a first segmented summary text; wherein the first segmented summary text is obtained by performing a text summary on the first segmented text; and treating the first segmented text and the first segmented summary text as text blocks.
[0131] In one embodiment, the text segmentation module 610 is further configured to input summary prompt words into the large language model; wherein the summary prompt words include the first segmented text and summary requirement information; and obtain a first summary text output by the large language model based on the summary prompt words; wherein the first summary text is the text in the first segmented text that meets the summary requirement information.
[0132] In one embodiment, the text segmentation module 610 is further used to input semantic segmentation prompt words into the large language model; wherein the semantic segmentation prompt words include the first merged text and semantic segmentation information; obtain the first segmented text and the first segmentation summary text output by the large language model based on the semantic segmentation prompt words; wherein the first segmented text is the text in the first merged text that meets the semantic segmentation information.
[0133] In one embodiment, the text filtering module 620 is further configured to input a filtering prompt word into the large language model; wherein the filtering prompt word includes multiple text blocks and filtering requirement information; and obtain a target text block output by the large language model based on the filtering prompt word; wherein the target text block is a text block among the multiple text blocks that meets the filtering requirement information.
[0134] In one embodiment, the feature extraction module 630 is further used to input feature extraction prompt words into the large language model; wherein the feature extraction prompt words include the target text block and feature requirement information; obtain feature attributes output by the large language model based on the feature extraction prompt words; wherein the feature attributes are text in the target text block that meets the feature requirement information.
[0135] In one embodiment, the text summary module 650 is further used to input text summary prompt words into the large language model; wherein the text summary prompt words include characteristic text blocks and text summary information; and obtain the acceptance and delivery standards output by the large language model based on the text summary prompt words; wherein the acceptance and delivery standards are the text in the characteristic text blocks that meets the text summary information.
[0136] Each module in the aforementioned receiving and delivery standard acquisition device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0137] In one embodiment, a computer device is provided, which may be a terminal. Its internal structure diagram may be as shown in Figure 11. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, where the wireless means may be implemented via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for obtaining receiving and sending standards. The display unit of the computer device is used to produce a visually visible image and may be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0138] Those skilled in the art will understand that the structure shown in FIG11 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.
[0139] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0140] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the processor executes the computer program, the steps in the above-mentioned method embodiments are implemented.
[0141] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0142] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0143] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.
Claims
1. A method for obtaining collection and delivery standards, characterized in that: The method comprises: Perform text segmentation on the target text to obtain multiple text blocks; Filtering out a target text block from the plurality of text blocks; wherein the target text block is a text block related to the collection and delivery standards; Performing feature extraction on the target text block to obtain feature attributes of the item; splicing the target text blocks with the same characteristic attributes to obtain characteristic text blocks; The characteristic text blocks are summarized to obtain the acceptance and delivery standards of the items.
2. The method for obtaining the collection and delivery standards according to claim 1, characterized in that: The step of segmenting the target text to obtain multiple text blocks includes: Segmenting the target text based on a preset number of words and / or preset symbols to obtain a plurality of segmented texts; The plurality of segmented text blocks are recursively segmented based on semantics to obtain a plurality of the text blocks.
3. The method for obtaining the collection and delivery standards according to claim 2, characterized in that: The plurality of segmented texts include: a first segmented text and a second segmented text, wherein the second segmented text is adjacent to the first segmented text; The recursive segmentation of the plurality of segmented text blocks based on semantics to obtain the plurality of text blocks includes: Performing text summarization on the first segmented text to obtain a first summary text; Merging the first segmented text, the second segmented text, and the first summary text to obtain a first merged text; According to the first summary text, semantic segmentation is performed on the combination of the first segmented text and the second segmented text in the first merged text to obtain a first segmented text and a second segmented text; a first segmented summary text is determined; wherein the first segmented summary text is obtained by performing textual summary on the first segmented text; The first segmented text and the first segmented summary text are taken as one text block.
4. The method for obtaining the collection and delivery standards according to claim 3, characterized in that: After semantic segmentation is performed on the combination of the first segmented text and the second segmented text in the first merged text according to the first summary text to obtain a first segmented text and a second segmented text, the method further includes: Determining a second segmentation summary text; wherein the second segmentation summary text is obtained by summarizing the second segmentation text and its corresponding context information; Using the second segmented summary text as a new first summary text; The original second segmented text and the third segmented text are used as the new first segmented text and the new second segmented text respectively; the third segmented text is adjacent to the original second segmented text; Return to the step of merging the first segmented text, the second segmented text, and the first summary text to obtain a first merged text, thereby obtaining a plurality of the text blocks.
5. The method for obtaining the collection and delivery standards according to claim 3, characterized in that: The step of summarizing the first segmented text to obtain a first summary text includes: Inputting summary prompt words into the large language model; wherein the summary prompt words include the first segmented text and summary requirement information; the summary requirement information is used to instruct the large language model to summarize the first segmented text; The first summary text output by the large language model according to the summary prompt words is obtained.
6. The method for obtaining the collection and delivery standards according to claim 4, characterized in that: The step of performing semantic segmentation on the combination of the first segmented text and the second segmented text in the first merged text according to the first summary text to obtain a first segmented text and a second segmented text includes: Inputting semantic segmentation prompt words into the large language model; wherein the semantic segmentation prompt words include the first merged text and semantic segmentation information; the semantic segmentation information is used to indicate that based on understanding the first summary text, the first segmented text, and the second segmented text, the combination of the first segmented text and the second segmented text is segmented, and the two segments of text obtained by segmentation are summarized; The first segmented text, the second segmented text, the first segmented summary text, and the second segmented summary text output by the large language model according to the semantic segmentation prompt words are obtained.
7. The method for obtaining the collection and delivery standards according to any one of claims 1 to 6, characterized in that: The step of selecting a target text block from the plurality of text blocks comprises: Inputting a screening prompt word into the large language model; wherein the screening prompt word includes the plurality of text blocks and screening requirement information; the screening requirement information includes output requirement information and output condition information; the output requirement information is used to indicate content information of an output result of the large language model; the output condition information is used to indicate condition information corresponding to whether the large language model outputs a yes or no result; The target text block is determined based on the output of the large language model according to the screening prompt words.
8. The method for obtaining the collection and delivery standards according to claim 7, characterized in that: The output requirement information includes: Determine whether the text is related to cross-border collection and delivery, and whether it involves specific requirements for the collection and delivery items or the collection and delivery entities, and give a result on whether the text content meets the requirements of the task, indicating the result with yes or no.
9. The method for obtaining the collection and delivery standards according to claim 7, characterized in that: The output condition information includes: The first output condition information is used to indicate that if the text content meets the first preset condition, the determination result is "yes"; The second output condition information is used to indicate that if the text content meets the second preset condition, the determination result is "no".
10. The method for obtaining the collection and delivery standards according to claim 7, wherein: The step of determining the target text block based on the output of the screening prompt word based on the large language model includes: The text block containing "yes" in the output result of the large language model is determined as the target text block.
11. The method for obtaining the collection and delivery standards according to any one of claims 1 to 6, characterized in that: The step of extracting features from the target text block to obtain feature attributes of the item includes: Inputting a feature extraction prompt word into the large language model; wherein the feature extraction prompt word includes the target text block and feature requirement information; the feature requirement information includes: extraction requirement information and association judgment information; the extraction requirement information is used to indicate the extraction requirement of the large language model; the association judgment information is used to instruct the large language model to determine whether the current text is associated with a preset item; The feature attribute corresponding to each target text block output by the large language model according to the feature extraction prompt word is obtained.
12. The method for obtaining the collection and delivery standards according to any one of claims 1 to 6, characterized in that: The step of summarizing the characteristic text block to obtain the acceptance and delivery standards of the item includes: Inputting a text summary prompt word into the large language model; wherein the text summary prompt word includes the characteristic text block and text summary information; the text summary information is used to instruct the large language model to extract the cross-border collection and delivery requirements of a certain item from the text, as well as the constraints that need to be taken into account when extracting the requirements; The receiving and sending standards output by the large language model according to the text summary prompt words are obtained.
13. The method for obtaining the collection and delivery standards according to any one of claims 1 to 6, characterized in that: The acceptance and delivery standards include: whether the item is allowed to be sent across borders, and / or the material information required during the acceptance and delivery process; The method further comprises: The collection and delivery standards of all items included in the target text are combined into a collection and delivery standard library.
14. A device for obtaining receiving and sending standards, characterized in that: The device comprises: The text segmentation module is used to segment the target text into multiple text blocks; A text screening module, configured to screen out a target text block from the plurality of text blocks; wherein the target text block is a text block related to the collection and delivery standards; A feature extraction module is used to extract features from the target text block to obtain feature attributes of the object; A text splicing module, configured to splice the target text blocks with the same characteristic attributes to obtain characteristic text blocks; The text summarizing module is used to perform text summarizing on the characteristic text blocks to obtain the acceptance and delivery standards of the items.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 13 are implemented.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
Citation Information
Patent Citations
Method for producing a document summary
CA2566013A1
Text partitioning method and device, computer equipment and storage medium
CN112733545A
Text abstract generation method and device, electronic equipment and readable medium
CN113673215A
Abstract generation method and device
CN115221311A
Long text abstract generation method based on hierarchical BERT model and label migration
CN116501861A