Big model-based publishing method and device, intelligent agent, equipment, medium and product
Through the category recommendation model and information generation model, product categories and information are automatically recommended and generated, which solves the problem of low efficiency in product creation in e-commerce platforms and realizes an efficient and accurate product release process.
Patent Information
- Application Number
- CN202510927401.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-10
Smart Images

Figure CN120765348A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical field of deep learning, large model, etc., which can be used in generative search, intelligent assistant, intelligent e-commerce, etc. application scenarios, more specifically, a large model-based publishing method, device, agent, equipment, medium and product. BACKGROUND
[0002] With the rapid development of the e-commerce field, as the core cornerstone of the e-commerce platform, the efficiency of creating and managing goods directly affects the operation effect and user experience of the platform. SUMMARY
[0003] The present disclosure provides a large model-based publishing method, device, agent, equipment, medium and product.
[0004] According to one aspect of the present disclosure, a large model-based publishing method is provided, comprising: in response to a publishing instruction, inputting multi-modal information of an item to be published obtained via an interactive interface into a category recommendation large model to obtain a target category to which the item belongs, wherein the multi-modal information includes at least one of a first image and an item description; in response to receiving at least one second image of the item, inputting the target category and the at least one second image into an information generation large model to display item information of the item on the interactive interface, wherein the first image and the second image are images of the item under different viewing angles; and in response to the item information passing verification, publishing the item based on the item information.
[0005] According to another aspect of the present disclosure, a large model-based publishing device is provided, comprising: a category recommendation module configured to, in response to a publishing instruction, input multi-modal information of an item to be published obtained via an interactive interface into a category recommendation large model to obtain a target category to which the item belongs, wherein the multi-modal information includes at least one of a first image and an item description; an information generation module configured to, in response to receiving at least one second image of the item, input the target category and the at least one second image into an information generation large model to display item information of the item on the interactive interface, wherein the first image and the second image are images of the item under different viewing angles; and a publishing module configured to, in response to the item information passing verification, publish the item based on the item information.
[0006] According to another aspect of the present disclosure, an artificial intelligence agent is provided, configured to execute the above method.
[0007] According to another aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the method.
[0008] According to another aspect of the present disclosure, a computer readable storage medium is provided, having stored thereon a computer program or instructions, which, when executed by a processor, implement the steps of the method.
[0009] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program or instructions, which, when executed by a processor, implement the steps of the method.
[0010] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0011] The above and other objects, features and advantages of the present disclosure will become more apparent from the following description when taken in conjunction with the accompanying drawings, in which:
[0012] Figure 1 A system architecture to which a large model-based publishing method can be applied is schematically shown according to an embodiment of the present disclosure;
[0013] Figure 2 A flowchart of a large model-based publishing method is schematically shown according to an embodiment of the present disclosure;
[0014] Figure 3A An example schematic diagram of a training process of a category recommendation large model is schematically shown according to an embodiment of the present disclosure;
[0015] Figure 3B An example schematic diagram of a process of inputting multi-modal information of an item to be published, acquired via an interactive interface, into a category recommendation large model to obtain a target category to which the item belongs is schematically shown according to an embodiment of the present disclosure;
[0016] Figure 4 An example schematic diagram of a category recommendation link interactive interface is schematically shown according to an embodiment of the present disclosure;
[0017] Figure 5A An example schematic diagram of a process of inputting a target category and at least one second image into an information generation large model to obtain item information of an item is schematically shown according to an embodiment of the present disclosure;
[0018] Figure 5BSchematically illustrates an example of a process of recognizing a second image and obtaining structured information according to an embodiment of the present disclosure;
[0019] Figure 6 Schematically shows an example schematic diagram of the interactive interface of the information generation phase according to an embodiment of the present disclosure;
[0020] Figure 7A Schematically illustrates an example schematic diagram of verifying item information based on the first method according to an embodiment of the present disclosure;
[0021] Figure 7B An example schematic diagram of an interactive interface displaying a first verification result according to an embodiment of the present disclosure is schematically shown;
[0022] Figure 8A Schematically illustrates an example schematic diagram of verifying item information based on the second method according to an embodiment of the present disclosure;
[0023] Figure 8B An example schematic diagram of an interactive interface displaying a second verification result according to an embodiment of the present disclosure is schematically shown;
[0024] Figure 9 An example schematic diagram of a publishing process based on a large model according to an embodiment of the present disclosure is schematically shown;
[0025] Figure 10 A block diagram of a publishing device based on a large model according to an embodiment of the present disclosure is schematically shown;
[0026] Figure 11 Schematically shows a structural block diagram of an intelligent agent of a large model according to an embodiment of the present disclosure; and
[0027] Figure 12 A block diagram of an electronic device suitable for implementing a large model-based publishing method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0028] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.
[0029] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0030] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0031] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0032] When creating a product, merchants typically need to manually fill out more than 20 fields. These fields are numerous and complex, requiring merchants to possess a high level of professional knowledge and experience. Furthermore, due to the stringent and complex review standards, product information submitted by merchants is often rejected for not meeting the requirements, increasing operational costs and time.
[0033] Furthermore, because the entire product creation process relies heavily on the merchant's personal skills, merchants must independently control the effectiveness and accuracy of the information they fill in. Due to the platform's weak automatic generation and verification capabilities, merchants are prone to multiple rejections due to incomplete or incorrectly formatted information, resulting in inefficient product releases. Furthermore, due to deficiencies in the platform's automated assisted entry and intelligent verification, it fails to effectively reduce merchants' input costs and fails to proactively identify and alert potential issues, resulting in a cumbersome and inefficient product release process.
[0034] To this end, embodiments of the present disclosure propose a macromodel-based publishing solution. For example, in response to a publishing instruction, multimodal information of an item to be published, obtained via an interactive interface, is input into a category recommendation macromodel to obtain a target category to which the item belongs, wherein the multimodal information includes at least one of a first image and an item description. In response to receiving at least one second image of the item, the target category and the at least one second image are input into a macromodel to display the item information on the interactive interface, wherein the first image and the second image are images of the item from different perspectives. In response to the item information passing verification, the item is published based on the item information.
[0035] According to the embodiments of the present disclosure, by utilizing a large category recommendation model to analyze the multimodal information of an item to be published to automatically recommend the target category to which the item belongs, the operational complexity of manually selecting a category by the subject is reduced, the possibility of problems with category selection is reduced, and the degree of automation and accuracy of category recommendation is improved. By utilizing a large information generation model to analyze the target category and the second image of the item to automatically generate the item information, the operational complexity of manually filling in information by the subject is reduced, the possibility of problems with information filling in is reduced, and the degree of automation and accuracy of item information generation is improved. On this basis, by verifying the item information and, if the verification passes, publishing the item based on the item information, the risk of erroneous publishing is reduced and the efficiency and quality of item publishing are improved.
[0036] In the technical solution of the present invention, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0037] In the technical solution of the present invention, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.
[0038] Figure 1 The schematic diagram shows the system architecture to which the large model-based publishing method according to the embodiment of the present disclosure can be applied. Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.
[0039] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104 and a server 105. The network 104 is used as a medium for providing a communication link between different devices.
[0040] It should be noted that the publishing method based on the large model provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the publishing device based on the large model provided in the embodiment of the present disclosure can generally be set in the server 105.
[0041] Alternatively, the large model-based publishing method provided in the embodiment of the present disclosure may also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the large model-based publishing apparatus provided in the embodiment of the present disclosure may also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0043] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0044] Figure 2 The flowchart of the publishing method based on the large model according to the embodiment of the present disclosure is schematically shown.
[0045] like Figure 2 As shown, the large model-based publishing method 200 includes operations S210 to S230.
[0046] In operation S210, in response to a publishing instruction, multimodal information of the item to be published obtained through the interactive interface is input into a category recommendation macro model to obtain a target category to which the item belongs, wherein the multimodal information includes at least one of a first image and an item description.
[0047] In operation S220, in response to receiving at least one second image of the item, a target category and at least one second image input information are input to generate a large model to display item information of the item on an interactive interface, wherein the first image and the second image are images of the item at different viewing angles.
[0048] In operation S230 , in response to the item information passing the verification, the item is released based on the item information.
[0049] A release instruction may refer to an instruction for initiating the item release process. The triggering method for a release instruction can be configured based on actual business needs and is not limited here. For example, a release instruction may be triggered by a subject clicking a "Release Item" button on an interactive interface. Alternatively, a release instruction may be triggered by a subject inputting at least one of a first image and a description of the item to be released on the interactive interface. Alternatively, a release instruction may be triggered by a subject uttering the words "I want to release an item" via voice. In one example, a subject may refer to a merchant, and an item may refer to a commodity.
[0050] After detecting a publishing instruction, multimodal information about the item to be published can be obtained via the interactive interface. Multimodal information can include information of at least two data types, such as visual, textual, and audio data types. In one example, the multimodal information can include a visual first image and a textual item description. In another example, the multimodal information can include only the visual first image. In this case, the item description can be obtained by analyzing the first image using a large model. The item description can refer to the item name.
[0051] After obtaining the multimodal information, the multimodal information can be input into the category recommendation model so that the category recommendation model determines the target category to which the item belongs by analyzing the first image and the item description of the item. The category recommendation model can be pre-trained. The target category refers to the category to which the item belongs. For example, if the item included in the first image is a smartphone, the target category can be "electronic products > mobile phones". It should be noted that the above-mentioned category recommendation model can introduce a feedback mechanism to dynamically optimize the category recommendation model based on each published item and the target category to which the item belongs.
[0052] In order to automatically generate item information for an item using the information generation macromodel, at least one second image of the item may be further acquired. The first image and the second image are images of the item from different perspectives. The perspectives may be determined based on the distance and angle of the item relative to the image capture device. For example, the first image may capture the main body of the item, while the second image may capture information about the back of the item, or the second image may capture information about the front of the item.
[0053] After obtaining the at least one second image, the target category and the at least one second image can be input into an information generation model, so that the information generation model generates item information for the item by analyzing the target category and the second image. The information generation model can be pre-trained. Item information can refer to attribute information of the item, which corresponds to the target category to which the item belongs.
[0054] After obtaining the item information, the generated item information can be displayed to the merchant for confirmation or modification. It should be noted that the above information generation model can introduce a feedback mechanism to dynamically optimize the information generation model based on each published item and the item information of the item.
[0055] After obtaining the item information, the item information can be verified to obtain a verification result indicating whether the item information has passed the verification. The specific verification method can be configured according to actual business needs and is not limited here. For example, the item information can be automatically verified using a risk verification model; alternatively, verification rules can be set and the item information can be automatically verified based on the verification rules; alternatively, a manual verification process can be introduced to allow the user to further verify the item information. If the verification result indicates that the item information has passed the verification, the item can be released based on the item information; if the verification result indicates that the item information has failed the verification, the release of the item can be blocked.
[0056] According to the embodiments of the present disclosure, by utilizing a large category recommendation model to analyze the multimodal information of an item to be published to automatically recommend the target category to which the item belongs, the operational complexity of manually selecting a category by the subject is reduced, the possibility of problems with category selection is reduced, and the degree of automation and accuracy of category recommendation is improved. By utilizing a large information generation model to analyze the target category and the second image of the item to automatically generate the item information, the operational complexity of manually filling in information by the subject is reduced, the possibility of problems with information filling in is reduced, and the degree of automation and accuracy of item information generation is improved. On this basis, by verifying the item information and, if the verification passes, publishing the item based on the item information, the risk of erroneous publishing is reduced and the efficiency and quality of item publishing are improved.
[0057] Figure 3A An example diagram schematically illustrates the training process of a large category recommendation model according to an embodiment of the present disclosure.
[0058] like Figure 3A As shown in 300A, historical item release records 301 may be extracted and organized to obtain multiple sample mappings 302. Sample mappings 302 may include a correspondence between historical multimodal information 302_2 and historical categories 302_1. Historical multimodal information 302_2 may include at least one of historical images and historical item descriptions.
[0059] After obtaining the plurality of sample mappings 302 , the plurality of sample mappings 302 may be used to train the category recommendation large model 303 to obtain a trained category recommendation large model.
[0060] For each historical category 302_1, the number of sample mappings 302 corresponding to the historical category 302_1 may be greater than or equal to a preset number, so that the trained category recommendation model can fully learn the characteristics of the multimodal information corresponding to the historical category 302_1, thereby being able to perform targeted category recommendations based on at least one of the new first image and the item description. For example, the preset number may be 50.
[0061] In one example, historical item release records 301 can be dynamically updated at predetermined intervals. Accordingly, sample map 302 can also be dynamically updated at predetermined intervals. For example, after each item release, the multimodal information used to execute the release and the target category to which the item belongs can be stored in a cache. At predetermined intervals, the stored information can be retrieved from the cache and added to historical item release records 301, thereby dynamically updating sample map 302 and optimizing the category recommendation model 303. For example, the predetermined interval can be daily or weekly.
[0062] In another example, historical item release records 301 may also be updated in response to a predetermined event being triggered. Predetermined events can be configured based on actual business needs and are not limited here. For example, a predetermined event may be a promotional activity. In this case, information corresponding to the category involved in the promotional activity can be retrieved from historical item release records 301, and the sample mapping 302 corresponding to the category can be updated.
[0063] According to the embodiments of the present disclosure, by training a large category recommendation model using sample mappings that characterize the correspondence between historical multimodal information and historical categories, the comprehensiveness and reliability of category recommendations are improved. Furthermore, regular updates to sample mappings ensure that the model adapts promptly to changes in published items, helping to improve the timeliness and accuracy of recommended target categories.
[0064] Figure 3B The diagram schematically shows an example process of inputting multimodal information of an item to be published obtained through an interactive interface into a category recommendation model to obtain a target category to which the item belongs according to an embodiment of the present disclosure.
[0065] like Figure 3B As shown, in 300B, after obtaining the multimodal information 304 of the item to be released, at least one of the first image 304_1 and the item description 304_2 included in the multimodal information 304 can be input into the category recommendation model 303, and at least one candidate category 305 can be output and displayed on the interactive interface.
[0066] After receiving the multimodal information 304, the category recommendation model 303 can determine the relevance of each predetermined category with the first image and the item description, obtain the relevance of each of the multiple predetermined categories, and sort the multiple predetermined categories in descending order of relevance to obtain a category sequence.
[0067] The relevance may refer to a quantified similarity value between the image features of the first image 304_1 , the text features of the item description 304_2 , and the predetermined category, and may be represented, for example, by cosine similarity, cross-attention weight, or probability distribution value.
[0068] On this basis, the predetermined categories that are ranked first in the category sequence can be determined as candidate categories 305. The predetermined ranking can be configured based on actual business needs and is not limited here. For example, the predetermined ranking can be less than or equal to 3. Alternatively, the predetermined ranking can be dynamically adjusted based on the degree of relevance. For example, the predetermined ranking can be determined based on the degree of relevance being greater than a predetermined relevance threshold.
[0069] After the interactive interface displays at least one candidate category 305, the subject can select a target category 306 from the at least one candidate category 305 displayed on the interactive interface. Alternatively, if the subject still does not perform a selection operation after a second predetermined period of time, the candidate category 305 with the highest degree of relevance to the first image 304_1 and the item description 304_2 among the at least one candidate category 305 can be selected as the target category 306.
[0070] According to the embodiments of the present disclosure, category recommendations for items are implemented by analyzing at least one of the first image and the item description using a large category recommendation model, thereby improving the automation and accuracy of category recommendations. Furthermore, in the interactive method for determining target categories, the visual display and ranking of candidate categories allows users to quickly confirm or select appropriate categories, improving interaction efficiency. This ranking-based method simplifies the publishing process.
[0071] Figure 4 An example diagram of an interactive interface for a category recommendation process according to an embodiment of the present disclosure is schematically shown.
[0072] like Figure 4 As shown, in 400, after obtaining the multimodal information of the item to be published, the first image and the item description included in the multimodal information can be input into the category recommendation model to obtain at least one candidate category, and the at least one candidate category is displayed on the interactive interface.
[0073] For example, the at least one candidate category may include category 1, category 2, and category 3. The subject may select a target category (eg, category 2) from the at least one candidate category through a selection operation on the interactive interface.
[0074] Figure 5A The diagram schematically illustrates an example process of generating a large model from a target category and at least one second image input information to obtain item information of an item according to an embodiment of the present disclosure.
[0075] like Figure 5A As shown, in 500A, after receiving the second image 501, the second image 501 can be recognized to obtain structured information 502. Based on this, the structured information 502 and at least one attribute field 504 for the item determined based on the target category 503 can be input into the information generation macro model 505, so that the information generation macro model 505 extracts information from the structured information 502 based on the attribute field 504 to obtain item information 506 including the attribute values of the at least one attribute field 504.
[0076] Structured information 502 may refer to information extracted from second image 501 and represented in a structured form. For example, structured information 502 identified from second image 501 may include "Origin: City A" and "Shelf Life: 12 Months." The method for obtaining structured information 502 can be configured based on actual business needs and is not limited here. For example, structured information 502 may be obtained by recognizing second image 501 using optical character recognition (OCR) technology. Alternatively, structured information 502 may be obtained by recognizing second image 501 using object detection technology.
[0077] Attribute field 504 can refer to a specific attribute name in the item information, used to describe the characteristics of the item. Examples include "place of origin" and "shelf life." Attribute value can refer to the value corresponding to the attribute field. For example, "City A" can be the attribute value for the "place of origin" attribute field, and "12 months" can be the attribute value for the "shelf life" attribute field. In one example, a prompt template can be used to guide the information generation macro model 505 in extracting information from the structured information 502 based on the attribute field 504. For example, the prompt template can be "Extract content corresponding to the attribute field from the structured information."
[0078] It should be noted that items belonging to different target categories have different attribute fields corresponding to them, so the prompt template can be determined based on the attribute fields corresponding to the target category. For example, taking the target category of "clothing" as an example, the attribute fields may include material, size, etc., so the prompt template can be "extract content corresponding to material and size from the structured information." Alternatively, taking the target category of "food" as an example, the attribute fields may include origin and expiration date, etc., so the prompt template can be "extract content corresponding to origin and expiration date from the structured information."
[0079] According to an embodiment of the present disclosure, by utilizing the structured information extracted from the second image and at least one attribute field for the item determined based on the target category, a large model can be generated based on the attribute field guidance information to obtain the corresponding attribute value from the structured information, thereby reducing the complexity of manual input operations by the user, helping to improve the efficiency of item publishing, and avoiding the possible erroneous input problems that may exist in manual input, helping to improve the accuracy of item information.
[0080] Figure 5B An example schematic diagram of a process of recognizing a second image and obtaining structured information according to an embodiment of the present disclosure is schematically shown.
[0081] like Figure 5B As shown, in 500B, taking the target category of the item as food as an example, after receiving the second image 507 of the item, since the text content in the second image 507 is unstructured information, the second image 507 can be analyzed to extract the unstructured text content in the second image 507 and generate structured information of the second image 507.
[0082] For example, text detection can be performed on the second image 507 to obtain detection information, which can include category information and location information of each of the multiple text regions. For example, after performing text detection on the second image 507, text region 508_1, text region 508_2, text region 508_3, and text region 508_4 can be obtained.
[0083] The category information can represent the category of the text content included in the text area. The category information can include at least one of the following: attribute field category or attribute value category. The attribute field category can represent that the text content included in the text area belongs to the attribute field. The attribute value category can represent that the text content included in the text area belongs to the attribute value. For example, if the text content included in text area 508_1 is "Place of Origin:", the category information of the text area is the attribute field category. Alternatively, if the text content included in text area 508_2 is "Shelf Life:", the category information of the text area is the attribute value category.
[0084] The position information can represent the location of the text area. The position information can be used as a basis for extracting the area image corresponding to the text area from the second image 507. For example, the position information can be represented by a text detection frame. The text detection frame can include four corner points, that is, the position information can be represented by four coordinates.
[0085] After obtaining the detection information, the regional image of each text region can be cut out from the second image 507 according to the position information of each text region. Specific cutting methods may include at least one of the following: a threshold-based image segmentation method, a region-based image segmentation method, an edge-based image segmentation method, an image segmentation method based on a specific theory, an image segmentation method based on genetic coding, an image segmentation method based on wavelet transform, and an image segmentation method based on a neural network.
[0086] For example, based on the position information of the text area 508_1, the area image 509_1 can be cut out from the second image 507; based on the position information of the text area 508_2, the area image 509_2 can be cut out from the second image 507; based on the position information of the text area 508_3, the area image 509_3 can be cut out from the second image 507; based on the position information of the text area 508_4, the area image 509_4 can be cut out from the second image 507.
[0087] After obtaining the region images of the multiple text regions, text recognition can be performed on each region image to obtain recognition information. For example, text recognition can be performed on region image 509_1 to obtain text information 510_1; text recognition can be performed on region image 509_2 to obtain text information 510_2; text recognition can be performed on region image 509_3 to obtain text information 510_3; and text recognition can be performed on region image 509_4 to obtain text information 510_4.
[0088] After obtaining the identification information, the semantic features of each text information can be extracted, and semantic relationship information representing the semantic relationship between the multiple text information can be determined based on the semantic features. The method for obtaining the semantic relationship information can be configured according to actual business needs and is not limited here. For example, the semantic relationship information can be obtained by performing global feature extraction on the identification information. Alternatively, the semantic relationship information can also be determined based on auxiliary information and the identification information, and the auxiliary information can include at least one of the following: position information and a first feature map obtained by performing feature extraction on the second image 507.
[0089] By performing feature extraction on the second image 507, a second feature map at at least one scale is obtained. A first feature map is obtained based on the second feature map at at least one scale. The scale may refer to image resolution. Each scale may have at least one second feature map corresponding to the scale.
[0090] In one example, when the auxiliary information includes a first feature map, semantic relationship information can be determined based on the first feature map and the identification information. For example, feature extraction can be performed on the text information of multiple regional images to obtain a second feature map corresponding to the identification information. The first feature map and the second feature map corresponding to the identification information are fused to obtain a fused feature map. Semantic relationship information is determined based on the fused feature map.
[0091] In another example, when the auxiliary information includes a first feature map and location information, semantic relationship information can be determined based on the first feature map, location information, and identification information. For example, feature extraction can be performed on the text information of multiple regional images to obtain a second feature map corresponding to the identification information. The first feature map and the second feature map corresponding to the identification information are fused to obtain a fused feature map. Semantic relationship information is determined based on the fused feature map and location information.
[0092] According to an embodiment of the present disclosure, since the semantic relationship information is determined based on auxiliary information and identification information, the auxiliary information includes at least one of the second feature map and position information. By utilizing the auxiliary information and identification information to determine the semantic relationship information, the accuracy of the semantic relationship information is improved.
[0093] After obtaining the semantic relationship information, structured information can be generated based on the category information, identification information, and semantic relationship information determined based on the identification information. The structured information can include values corresponding to attribute field categories and values corresponding to attribute value categories. For example, the semantic relationship information can include the semantic relationship between text information 510_1 and text information 510_3, and the semantic relationship between text information 510_2 and text information 510_4. This can generate structured information 511_1 as "Place of Origin: City A" and structured information 511_2 as "Shelf Life: 12 Months."
[0094] Figure 6 An example schematic diagram of an interactive interface for an information generation phase according to an embodiment of the present disclosure is schematically shown.
[0095] like Figure 6 As shown, in 600, after obtaining at least one second image of the item, the target category and the at least one second image input information can be used to generate a large model to display the item information of the item on the interactive interface.
[0096] For example, the item information can include "origin: City A; shelf life: 12 months". The object can modify the item information (e.g., modify the attribute value corresponding to the shelf life) through a modification operation on the interactive interface.
[0097] In one example, after obtaining the item information, the item information can be verified based on a first manner to obtain a first verification result corresponding to the first manner. The first manner is used for text dimension verification. The first verification result refers to the result after the first manner verification, and is used to represent whether the item information passes the verification in the text dimension. For example, the first verification result can be "pass" or "fail, reason: description incomplete".
[0098] The specific first manner can be configured according to actual business needs, which is not limited here. For example, the first manner can include at least one of the following: field integrity check, format compliance check, and semantic reasonableness check, etc.
[0099] The field integrity check refers to checking whether the attribute value corresponding to the mandatory attribute field exists. For example, if "price" is a mandatory attribute field, if there is an attribute value corresponding to the price in the item information, the item information passes the field integrity check.
[0100] The format compliance refers to checking whether the attribute value corresponding to each attribute field conforms to the standard format of the attribute field. For example, the standard format corresponding to the attribute field "date" is "YYYYMMDD", and if the attribute value corresponding to the attribute field is "20250101", the item information passes the format compliance check.
[0101] The semantic reasonableness check refers to checking whether the attribute value corresponding to each attribute field is within the expected attribute value range. For example, the attribute value corresponding to the attribute field "size" is checked to see if it is within the range of S, M, and L. If so, the item information passes the semantic reasonableness check.
[0102] In another example, after obtaining the item information, the item information can also be verified based on the first manner to obtain a second verification result corresponding to the second manner. The second manner is used for risk dimension verification. The second verification result refers to the result after the second manner verification, and is used to represent whether the item information passes the verification in the risk dimension. For example, the result can be "pass" or "fail, reason: involves a certain risk type of risk".
[0103] After obtaining the first verification result and the second verification result, if both verification results are "pass", it can be determined that the item information passes the verification, and in this case, "verification passed, item can be published" can be displayed through the interactive interface.
[0104] According to the embodiments of the present disclosure, the first method, which verifies the text dimension, ensures the integrity and standardization of item information, reducing item release failures due to incomplete information. The second method, which verifies the risk dimension, effectively identifies and prevents the release of risky content. By combining the multi-dimensional verification mechanisms of the first and second methods, the accuracy and security of item information can be ensured.
[0105] Figure 7A An example schematic diagram of verifying item information based on the first method according to an embodiment of the present disclosure is schematically shown.
[0106] like Figure 7A As shown, in 700A, taking the reference item information that has passed the verification of the first method as "attribute field K1: attribute value V1; attribute field K2: attribute value V2; attribute field K3: attribute value V3" as an example, an example of the situation of item information that has not passed the verification of the first method is given.
[0107] In one example, the attribute value of each of at least one attribute field can be verified according to a preset format rule. The preset format rule defines the text format of each attribute field. In this case, for each attribute field, the attribute value of the attribute field can be checked for regularity, that is, to determine whether the attribute value conforms to the text format corresponding to the attribute field. For example, the preset format rule defines that the attribute value corresponding to the attribute field K1 is plain text and has a length of no more than 10 characters. For item information 701, if the attribute value V4 is not plain text or is longer than 10 characters, it can be determined that the item information 701 has failed the first method of verification.
[0108] In another example, for each attribute field, the correlation between the attribute field and the attribute value can be verified. For example, for item information 702, if attribute value V5 is not correlated with attribute field K2, then item information 702 can be determined to have failed the first verification method. Specifically, taking attribute field K2 as "color," if attribute value V5 is "3," then attribute value V5 can be determined to be not correlated with attribute field K2.
[0109] In another example, the relevance of at least one attribute field to the target category can be verified. In this case, each attribute field can be checked for relevance to the target category. For example, for item information 703, if attribute field K6 is not relevant to the target category, it can be determined that item information 703 has failed the first verification method. Specifically, taking the target category as "food," if attribute field K6 is "size," it can be determined that attribute field K6 is not relevant to the target category.
[0110] It should be noted that risk thresholds can be set for each of the aforementioned first methods. For each type, if it is determined that the item information fails verification for that type of first method, the severity of the issue can be assessed to obtain an assessment value. The first verification result can then be adjusted based on the assessment value and the corresponding risk threshold. Specifically, if the assessment value exceeds the risk threshold, indicating a serious issue, no adjustment to the first verification result is required. If the assessment value does not exceed the risk threshold, indicating a relatively minor issue, the first verification result can be adjusted to indicate that the item information passes verification for the first method.
[0111] According to the embodiments of the present disclosure, by adopting a first method based on preset format rules to verify item information, it is possible to ensure that attribute values conform to a specific format, thereby reducing publication failures caused by format errors; by adopting a first method based on the degree of correlation between attribute fields and target categories to verify item information, it is possible to improve classification accuracy; by adopting a first method based on the degree of correlation between attribute fields and attribute values to verify item information, it is possible to ensure the semantic correctness of the item information; thereby, the text dimension of the item information can be fully verified, which helps to improve the accuracy of the item information.
[0112] Figure 7B An example schematic diagram of an interactive interface displaying a first verification result according to an embodiment of the present disclosure is schematically shown.
[0113] like Figure 7B As shown, in 700B, in response to at least one of the target field and the target value in the item information failing to pass verification, at least one of the target field and the target value can be displayed on the interactive interface according to a preset display method.
[0114] The preset display methods may include at least one of the following: highlighting at least one of the target fields and target values that have not passed the verification based on a preset color; marking at least one of the target fields and target values that have not passed the verification based on a preset shape; identifying at least one of the target fields and target values that have not passed the verification based on a preset identifier, etc.
[0115] In 700B, for example, if the target category of an item is "food," and the attribute values of at least one attribute field are verified according to the preset formatting rules, it can be determined that the "Shelf Life: 9.5 Months" in the item information fails verification. Therefore, "9.5" can be marked based on the preset shape. In this case, the object can modify "9.5" through the interactive interface to obtain the modified information "9," and use this modified information to replace "Shelf Life: 9.5 Months" in the item information with "Shelf Life: 9 Months."
[0116] When verifying the relevance of at least one attribute field to the target category, it may be determined that "Material: Material B" in the item information fails verification, and thus "Material: Material B" may be highlighted using a preset color. In this case, the subject may delete "Material: Material B" through the interactive interface.
[0117] When verifying the correlation between the attribute field and the attribute value, it can be determined that the "Flavor: Buy One Get One Free" in the item information fails verification. Therefore, the "Flavor: Buy One Get One Free" can be identified based on a preset identifier. In this case, the subject can modify the "Buy One Get One Free" option through the interactive interface to obtain the modified information "Spicy." This modified information can then be used to replace the "Flavor: Buy One Get One Free" in the item information with "Flavor: Spicy."
[0118] According to the embodiments of the present disclosure, since fields or values that fail verification can be automatically detected and prompted, the time and complexity of users manually finding errors are reduced; and through preset display methods and modification operations, users can quickly locate problems and correct them, further simplifying the publishing process; in addition, through interactive verification and modification mechanisms, the efficiency and accuracy of the item publishing process are improved.
[0119] Figure 8A An example schematic diagram of verifying item information based on the second method according to an embodiment of the present disclosure is schematically shown.
[0120] like Figure 8A As shown in Figure 800A, a risk verification macromodel for that risk type and multimodal information about the item to be verified can be determined based on the risk type. Risk types can include fraud, misuse, fictitious activities, and exaggerated advertising. The multimodal information to be verified can include at least one of an item image 801 and item text 802 corresponding to the risk type. Item image 801 can include a first image 801_1, a second image 801_2, an item details image 801_3, and an item specifications image 801_4. Item text 802 can include an item description 802_1, item information 802_2, item specifications 802_3, and a target category 802_4.
[0121] For example, for risk type 1, the item image 801 may include the first image 801_1, the item details image 801_3, and the item specifications image 801_4, and the item text 802 may include the item description 802_1 and the item information 802_2. Alternatively, for risk type 2, the item image 801 may include the first image 801_1, and the item text 802 may include the item description 802_1.
[0122] For another example, for risk type 3, a risk verification model 804 can be used to process text information 803 identified from item image 801. By inputting item image 801 into risk verification model 804, a second verification result 805 is obtained. Alternatively, for risk type 4, a risk verification model 806 can be used to process item image 801 and item text 802. By inputting item image 801 and item text 802 into risk verification model 806, a second verification result 807 is obtained.
[0123] After obtaining the second verification result 805 and the second verification result 807, if the second verification result 805 indicates that the item information has passed the verification of the second method, the item can be published based on the item information; if the second verification result 807 indicates that the item information has not passed the verification of the second method, the item can be published based on the item information.
[0124] If second verification result 805 indicates that the item information failed verification using the second method, operation S810 may be executed. If second verification result 807 indicates that the item information failed verification using the second method, operation S810 may be executed. In operation S810, it may be determined whether the accuracy rate is less than a preset accuracy threshold. The preset accuracy threshold can be used to determine the reliability of the verification result. The preset accuracy threshold can be configured based on actual business needs and is not limited here. For example, the preset accuracy threshold may be 99%.
[0125] If the accuracy is greater than or equal to the preset accuracy threshold, that is, the accuracy is greater than or equal to 99%, it means that the possibility of the item information being risky is greater than 99%, and thus operation S820 can be executed. In operation S820, the item release is intercepted.
[0126] If the accuracy rate is less than the preset accuracy rate threshold, that is, the accuracy rate is less than 99%, it means that the possibility that the item information has a risk is less than 99%, and thus risk warning information 808 indicating the risk type can be output to the object through the interactive interface.
[0127] According to the embodiments of the present disclosure, since the risk verification model used to perform risk verification on item information, as well as the specifically selected item images and item texts, all correspond to the risk types, the pertinence and comprehensiveness of risk identification can be improved, thereby helping to improve the accuracy of the second verification result.
[0128] Figure 8B An example schematic diagram of an interactive interface displaying a second verification result according to an embodiment of the present disclosure is schematically shown.
[0129] like Figure 8BAs shown, in 800B, after obtaining multimodal information for the risk type, at least one of the item image and item text included in the multimodal information can be input into the risk verification model to obtain a second verification result.
[0130] If the second verification result indicates that the item information has failed verification using the second method and the accuracy is less than a preset accuracy threshold, a risk warning message indicating the risk type may be displayed to the subject via the interactive interface. For example, the risk warning message may read "Type C risk exists, please perform operation D." The subject may re-edit the item information by clicking the "Edit Item" control on the interactive interface.
[0131] Figure 9 An example diagram of a publishing process based on a large model according to an embodiment of the present disclosure is schematically shown.
[0132] like Figure 9 As shown in 900, object 901 can initiate a publishing instruction for an item through an interactive interface. In response to the publishing instruction, multimodal information 902 of the item to be published obtained through the interactive interface can be input into a category recommendation model 903 to obtain a target category 904 to which the item belongs.
[0133] In response to receiving the at least one second image 905 of the item, the target category 904 and the at least one second image 905 may be input into information to generate a macro model 906 to display item information 907 of the item on an interactive interface.
[0134] After obtaining item information 907, item information 907 can be verified using the first and second methods described above, and operation S910 can be executed. In operation S910, does item information 907 pass verification? If not, the content requiring modification can be displayed to object 901 via the interactive interface. If so, operation S920 can be executed. In operation S920, does manual review pass? If not, the content requiring modification can be displayed to object 901 via the interactive interface. If so, operation S930 can be executed. In operation S930, the item can be published.
[0135] The above are merely exemplary embodiments, but are not limited thereto. Other large model-based publishing methods known in the art may also be included, as long as they can improve the efficiency and quality of article publishing.
[0136] Figure 10 A block diagram of a large model-based publishing device according to an embodiment of the present disclosure is schematically shown.
[0137] like Figure 10As shown, the publishing device 1000 based on the big model may include a category recommendation module 1010 , an information generation module 1020 and a publishing module 1030 .
[0138] The category recommendation module 1010 is used to input the multimodal information of the item to be published obtained through the interactive interface into the category recommendation model in response to the publishing instruction to obtain the target category to which the item belongs, wherein the multimodal information includes at least one of the first image and the item description.
[0139] The information generation module 1020 is used to generate a large model based on the target category and the at least one second image input information in response to receiving at least one second image of the item, so as to display the item information of the item on the interactive interface, wherein the first image and the second image are images of the item at different perspectives.
[0140] The publishing module 1030 is configured to publish the item based on the item information in response to the item information passing the verification.
[0141] According to an embodiment of the present disclosure, the category recommendation module 1010 may include a recommendation submodule, a first determination submodule, and a second determination submodule.
[0142] The recommendation submodule is used to input at least one of the first image and the item description into the category recommendation model to display at least one candidate category on the interactive interface, wherein the candidate category is a category whose relevance to the first image and the item description is ranked second in the top preset position.
[0143] a first determining submodule, configured to determine, in response to a selection operation on any candidate category among at least one candidate category, a candidate category as a target category;
[0144] The second determining submodule is configured to select the candidate category with the highest correlation with the first image and the item description among the at least one candidate category as the target category.
[0145] According to an embodiment of the present disclosure, the category recommendation model is obtained by training using sample mapping, which characterizes the correspondence between historical multimodal information and historical categories. The historical multimodal information includes at least one of historical images and historical object descriptions. The historical mapping is dynamically updated at predetermined intervals.
[0146] According to an embodiment of the present disclosure, item information includes an attribute value of each of at least one attribute field.
[0147] According to an embodiment of the present disclosure, the information generating module 1020 may include an identifying submodule and a generating submodule.
[0148] The recognition submodule is used to recognize the second image and obtain structured information.
[0149] The generating sub-module is configured to input the structured information and at least one attribute field for the item determined based on the target category into the information generation large model, so that the information generation large model performs information extraction from the structured information based on the attribute field, to obtain respective attribute values of the at least one attribute field.
[0150] According to an embodiment of the present disclosure, the identifying sub-module can include a text detection unit, a text recognition unit, and a generating unit.
[0151] The text detection unit is configured to perform text detection on the second image to obtain detection information, where the detection information includes category information and position information of each of a plurality of text regions.
[0152] The text recognition unit is configured to perform text recognition on respective region images of the plurality of text regions obtained based on the position information and the second image, to obtain recognition information, where the recognition information includes text information of each of the plurality of region images.
[0153] The generating unit is configured to generate structured information according to the category information, the recognition information, and semantic relationship information determined based on the recognition information, where the semantic relationship information includes semantic relationships between the plurality of text information.
[0154] According to an embodiment of the present disclosure, the semantic relationship information is determined according to auxiliary information and the recognition information, and the auxiliary information includes at least one of the position information and a first feature map obtained by performing feature extraction on the second image.
[0155] According to an embodiment of the present disclosure, in a case where the auxiliary information includes the first feature map, the semantic relationship information is obtained by: fusing the first feature map and a second feature map corresponding to the recognition information to obtain a fused feature map; and determining the semantic relationship information according to the fused feature map.
[0156] According to an embodiment of the present disclosure, in a case where the auxiliary information further includes the position information, the semantic relationship information is obtained by: determining the semantic relationship information according to the fused feature map and the position information.
[0157] According to an embodiment of the present disclosure, the publishing device 1000 based on the large model can further include a verification module and a determination module.
[0158] The verification module is configured to perform verification on the item information based on a first mode and a second mode respectively, to obtain a first verification result corresponding to the first mode and a second verification result corresponding to the second mode, where the first mode is used for performing text dimension verification, and the second mode is used for performing risk dimension verification.
[0159] The determination module is configured to determine that the item information passes the verification in response to the first verification result indicating that the item information passes the verification in the first manner and the second verification result indicating that the item information passes the verification in the second manner.
[0160] According to an embodiment of the present disclosure, item information includes an attribute value of each of at least one attribute field.
[0161] According to an embodiment of the present disclosure, the check module may include a first check submodule, a second check submodule, and a third check submodule.
[0162] The first verification submodule is configured to verify the attribute value of each of the at least one attribute fields according to a preset format rule, wherein the preset format rule defines the text format of each attribute field.
[0163] The second verification submodule is used to verify the relevance between at least one attribute field and the target category.
[0164] The third verification submodule is used to verify the correlation between each attribute field and the attribute value.
[0165] According to an embodiment of the present disclosure, the verification module may further include a display submodule and a replacement submodule.
[0166] The display submodule is configured to display at least one of the target field and the target value on the interactive interface according to a preset display mode in response to at least one of the target field and the target value failing to pass the verification in the item information.
[0167] The replacement submodule is configured to replace at least one of the target field and the target value with the modification information in response to a modification operation of the object on at least one of the target field and the target value.
[0168] According to an embodiment of the present disclosure, the second method is used to verify whether an item has a risk of a risk type.
[0169] According to an embodiment of the present disclosure, the verification module may include a third determination submodule and a fourth verification submodule.
[0170] The third determination submodule is configured to determine a risk verification macromodel for the risk type and multimodal information of the item to be verified, wherein the multimodal information to be verified includes at least one of an item image and an item text corresponding to the risk type.
[0171] The fourth verification submodule is used to input the multimodal information to be verified into the risk verification model to obtain a second verification result.
[0172] According to an embodiment of the present disclosure, the verification module may further include an interception submodule and an output submodule.
[0173] The interception submodule is configured to intercept the release of the item in a case where the second check result indicates that the item information fails to pass the check by the second mode and the accuracy is greater than or equal to the preset accuracy threshold.
[0174] The output submodule is configured to output, to the object through an interactive interface, risk prompt information in a case where the second check result indicates that the item information fails to pass the check by the second mode and the accuracy is less than the preset accuracy threshold, where the risk prompt information is used to indicate a risk type.
[0175] Any one or more of the modules, submodules, units, and sub-units according to embodiments of the present disclosure, or at least part of the functions of any one or more of the modules, submodules, units, and sub-units, can be implemented in one module. Any one or more of the modules, submodules, units, and sub-units according to embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules, submodules, units, and sub-units according to embodiments of the present disclosure can be implemented at least in part as a hardware circuit, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on a substrate, a system in a package, an application-specific integrated circuit (ASIC), or any other reasonable manner of hardware or firmware through integration or packaging of circuits, or in any one of software, hardware, and firmware or in a proper combination of any of the foregoing. Alternatively, one or more of the modules, submodules, units, and sub-units according to embodiments of the present disclosure can be implemented at least in part as computer program modules that can perform corresponding functions when the computer program modules are run.
[0176] It should be noted that the part of the release device based on the large model in the embodiments of the present disclosure corresponds to the part of the release method based on the large model in the embodiments of the present disclosure, and the description of the part of the release device based on the large model is specifically referred to the part of the release method based on the large model, which will not be repeated here.
[0177] Figure 11 A structural block diagram of an agent of a large model according to an embodiment of the present disclosure is schematically shown.
[0178] In embodiments of the present disclosure, inspired by the Von Neumann structure in modern computer theory, as shown in Figure 11 The AI agent 1100 can include five core modules: an input module 1110, a control module 1120, a storage module 1130, an operation module 1140, and an output module 1150.
[0179] Input module 1110 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or the external environment) and converting it into a format that AI agent 1100 can understand and process. Input module 1110 is the primary link for AI agent 1100 to interact with the outside world. It enables AI agent 1100 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0180] In an example, the input module 1110 may input the first image, the item description, and the second image described above.
[0181] The control module 1120 is the core support for the AI agent 1100 to handle complex tasks. The control module 1120 can execute the large model-based publishing method described above.
[0182] In the example, the control module 1120 will continuously interact with the storage module 1130, the computing module 1140, and / or the output module 1150 during operation. However, it should be noted that in the embodiment of the present disclosure, the control module 1120 acts as a single initiator to initiate communication with the storage module 1130, the computing module 1140, and / or the output module 1150, and there is no communication coupling between the storage module 1130, the computing module 1140, and the output module 1150.
[0183] In this example, the performance of control module 1120 may be closely related to the large model underlying AI agent 1100. To fully leverage the capabilities of the large language model, the internal structure of control module 1120 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.
[0184] The storage module 1130 can be responsible for memorizing candidate category mappings for different application scenarios and the multimodal big model applicable to the application scenario. The category recommendation big model, information generation big model, and risk identification big model mentioned above can be included in the storage module 1130.
[0185] In this example, after receiving the release instruction, AI agent 1100 can trigger the macromodel release process to release the item. During this process, AI agent 1100 can retrieve model information for the category recommendation macromodel, information generation macromodel, and risk identification macromodel from storage module 1130 and feed it back to control module 1120. Control module 1120 can then pass the fed-back model information for the category recommendation macromodel, information generation macromodel, and risk identification macromodel to output module 1150.
[0186] The operation module 1140 can be regarded as a predefined tool library, and the tools for text detection and text recognition as described above can be included in the operation module 1140 .
[0187] In the example, when the AI agent 1100 needs to process data, it can call the relevant tools for text detection and text recognition from the operation module 1140 and feed it back to the control module 1120. Then, the control module 1120 can use the tools for text detection and text recognition that are fed back to generate the structured information and pass the structured information to the output module 1150. It can be understood that although the large language model has excellent language understanding and generation capabilities, it is the same as a human being. Without the help of any tools, the tasks that can be solved are very limited. When the AI agent 1100 is given the ability to call tools, it can achieve tasks such as completing text detection with the help of tools for text detection and completing text recognition with the help of tools for text recognition.
[0188] The output module 1150 can output the target category and item information described above.
[0189] The AI agent 1100 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.
[0190] Figure 12 A block diagram of an electronic device suitable for implementing a large model-based publishing method according to an embodiment of the present disclosure is schematically shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0191] like Figure 12 As shown, device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. RAM 1203 may also store various programs and data required for the operation of device 1200. Computing unit 1201, ROM 1202, and RAM 1203 are connected to each other via a bus 1204. An input / output (I / O) interface 1205 is also connected to bus 1204.
[0192] Various components in device 1200 are connected to I / O interface 1205, including an input unit 1206, such as a keyboard and mouse; an output unit 1207, such as various types of displays and speakers; a storage unit 1208, such as a magnetic disk and optical disk; and a communication unit 1209, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1209 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0193] Computing unit 1201 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Computing unit 1201 performs the various methods and processes described above, such as the large-model-based publishing method. For example, in some embodiments, the large-model-based publishing method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by computing unit 1201, one or more steps of the large-model-based publishing method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to execute the large model-based publishing method in any other appropriate manner (for example, by means of firmware).
[0194] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0195] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0196] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0197] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0198] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0199] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0200] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0201] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A publishing method based on a large model, comprising: In response to the publishing instruction, inputting multimodal information of the item to be published obtained via the interactive interface into the category recommendation macromodel to obtain a target category to which the item belongs, wherein the multimodal information includes at least one of the first image and the item description; In response to receiving at least one second image of the item, generating a macro model using the target category and the at least one second image input information to display item information of the item on the interactive interface, wherein the first image and the second image are images of the item at different viewing angles; and In response to the item information passing the verification, the item is released based on the item information.
2. The method according to claim 1, wherein Inputting the multimodal information of the to-be-published item obtained through the interactive interface into the category recommendation macromodel to obtain the target category to which the item belongs includes: Inputting at least one of the first image and the item description into the category recommendation model to display at least one candidate category on the interactive interface, wherein the candidate category is a category that has a relevance to the first image and the item description that ranks second among the top preset categories; and In response to a selection operation on any one of the at least one candidate category, the candidate category is determined as the target category; or, the candidate category with the highest degree of relevance to the first image and the item description among the at least one candidate category is determined as the target category.
3. The method according to claim 2, wherein: The category recommendation model is trained using sample mapping, which represents the correspondence between historical multimodal information and historical categories. The historical multimodal information includes at least one of historical images and historical item descriptions. The sample mapping is dynamically updated at predetermined intervals.
4. The method according to any one of claims 1 to 3, wherein The item information includes at least one attribute value of each attribute field; Generating a large model from the target category and the at least one second image input information to display the item information of the item on the interactive interface includes: recognizing the second image to obtain structured information; as well as The structured information and at least one attribute field for the item determined based on the target category are input into the information generation model, so that the information generation model extracts information from the structured information based on the attribute field to obtain the attribute value of each of the at least one attribute field.
5. The method according to claim 4, wherein The identifying the second image to obtain structured information includes: performing text detection on the second image to obtain detection information, wherein the detection information includes category information and position information of each of the plurality of text regions; Performing text recognition on respective region images of a plurality of text regions obtained based on the position information and the second image to obtain recognition information, wherein the recognition information includes text information of the respective region images; and The structured information is generated according to the category information, the identification information, and semantic relationship information determined based on the identification information, wherein the semantic relationship information includes semantic relationships between the plurality of text information.
6. The method according to claim 5, wherein: The semantic relationship information is determined based on auxiliary information and the identification information, and the auxiliary information includes at least one of the following: the position information and a first feature map obtained by performing feature extraction on the second image.
7. The method according to claim 6, wherein: In the case where the auxiliary information includes the first feature map, the semantic relationship information is obtained in the following manner: Fusing the first feature map with a second feature map corresponding to the identification information to obtain a fused feature map; and The semantic relationship information is determined according to the fused feature map.
8. The method according to claim 7, wherein: In the case where the auxiliary information also includes the position information, the semantic relationship information is obtained in the following manner: The semantic relationship information is determined according to the fused feature map and the position information.
9. The method according to any one of claims 1 to 3, further comprising, before releasing the item based on the item information in response to the item information passing verification: The item information is verified based on a first method and a second method respectively, to obtain a first verification result corresponding to the first method and a second verification result corresponding to the second method, wherein: The first method is used to verify the text dimension, and the second method is used to verify the risk dimension; as well as In response to the first verification result indicating that the item information passes the first verification method and the second verification result indicating that the item information passes the second verification method, it is determined that the item information passes the verification.
10. The method according to claim 9, wherein: The item information includes at least one attribute value of each attribute field; Verifying the item information based on the first method includes at least one of the following: Verifying the attribute value of each of the at least one attribute field according to a preset format rule, wherein the preset format rule defines the text format of each of the attribute fields; Verifying the relevance of the at least one attribute field to the target category; and For each of the attribute fields, the correlation between the attribute field and the attribute value is verified.
11. The method according to claim 10, further comprising: In response to at least one of the target field and the target value in the item information failing to pass verification, displaying at least one of the target field and the target value on the interactive interface according to a preset display mode; as well as In response to a modification operation of the object on at least one of the target field and the target value, at least one of the target field and the target value is replaced with modification information.
12. The method according to claim 9, wherein The second method is used to verify whether the item has a risk of a risk type; Verifying the item information based on the second method to obtain the second verification result includes: Determining a risk verification macromodel for the risk type and multimodal information of the item to be verified, wherein the multimodal information to be verified includes at least one of an item image and an item text corresponding to the risk type; and The multimodal information to be verified is input into the risk verification model to obtain the second verification result.
13. The method according to claim 12, further comprising: If the second verification result indicates that the item information fails the verification in the second manner and the accuracy rate is greater than or equal to a preset accuracy rate threshold, blocking the release of the item; as well as When the second verification result indicates that the item information has not passed the verification of the second method and the accuracy is less than the preset accuracy threshold, risk warning information is output to the object through the interactive interface, wherein the risk warning information is used to indicate the risk type.
14. A publishing device based on a large model, comprising: a category recommendation module, configured to, in response to a publishing instruction, input multimodal information of an item to be published, obtained via the interactive interface, into a category recommendation macromodel to obtain a target category to which the item belongs, wherein the multimodal information includes at least one of a first image and an item description; an information generation module, configured to, in response to receiving at least one second image of the item, input the target category and the at least one second image into information to generate a macro model to display item information of the item on the interactive interface, wherein the first image and the second image are images of the item at different viewing angles; and A publishing module is configured to publish the item based on the item information in response to the item information passing verification.
15. An artificial intelligence agent configured to execute the method according to any one of claims 1 to 13.
16. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 13.
17. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
18. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.