Article signing processing method and device, storage medium and electronic equipment

By monitoring delivery scenario videos in the home environment and using intelligent big models to generate structured event records, the manual error and information island problems in traditional item sign-up processing are solved, efficient and accurate sign-up process management is achieved, and user experience is improved.

CN120492580APending Publication Date: 2025-08-15SHENZHEN QIHOO INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510575669.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

Traditional items sign-up and processing methods rely on manual verification and physical recording, and there are problems of manual operation errors, delayed feedback and information islands, making it difficult to achieve real-time monitoring and data integration, affecting the sign-up efficiency.

Method used

By monitoring item delivery videos in home scenarios, using intelligent big models to interact with delivery personnel, generating automatic interactive dialogue sets, and performing multi-modal data fusion, generating structured item sign-up scene description events and autonomous response records, and pushing them to the user.

Benefits of technology

It improves the efficiency and accuracy of sign-up processing, simplifies users' backtracking and verification of details of the sign-up process, and improves the security and user experience of the smart home system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492580A_ABST
    Figure CN120492580A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an article signing processing method and device, a storage medium and electronic equipment, and the method comprises the steps: monitoring an article delivery scene video in a home scene, and carrying out the article signing processing process with article delivery personnel through an intelligent large model in the article delivery scene, obtaining an automatic interactive dialogue set corresponding to the article signing processing process, and performing structured event generation based on the automatic interactive dialogue set and the article delivery scene video through an intelligent large model to obtain an article signing scene description event, and generating an article signing autonomous response record based on the article signing scene description event, and pushing the article signing autonomous response record to the user side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a method, device, storage medium, and electronic device for processing item receipt. Background Art

[0002] In recent years, delivery personnel have typically delivered packages to recipients, who then sign for them in person. This process relies on manual verification and physical record-keeping. While simple and feasible, this approach is subject to issues such as manual error, delayed feedback, and information silos, making it difficult to achieve real-time monitoring and data integration throughout the delivery process, thus impacting delivery efficiency. Summary of the Invention

[0003] The embodiments of this specification provide a method, device, storage medium, and electronic device for processing item receipt. The technical solution is as follows:

[0004] In a first aspect, an embodiment of this specification provides a method for processing item receipt, the method comprising:

[0005] Monitor the delivery scene video in the home scene, and use the intelligent big model to communicate with the delivery personnel during the delivery process;

[0006] Obtaining a set of automatic interactive dialogues corresponding to the item receipt processing process;

[0007] Generate an item receipt scene description event by performing structured event generation based on the automatic interactive dialogue set and the item delivery scene video through an intelligent big model, and generate an item receipt autonomous response record based on the item receipt scene description event;

[0008] Push the autonomous response record of the item receipt to the user end.

[0009] In a feasible implementation, the generating of structured events based on the automatic interactive dialogue set and the item delivery scene video by the intelligent big model to obtain the item receipt scene description event includes:

[0010] Perform semantic recognition on the automatic interactive dialogue set through an intelligent big model to determine multiple key sign-off nodes;

[0011] Based on the item delivery scene video and the automatic interactive dialogue set, the key sign-off nodes are mapped and aligned with multimodal node information using an intelligent big model to obtain multimodal key sign-off node features;

[0012] Based on the multimodal key receipt node features, an item receipt dialogue event template is used to perform structured event conversion processing to obtain an item receipt scene description event.

[0013] In a feasible implementation, based on the item delivery scene video and the automatic interactive dialogue set, the multimodal node information mapping and alignment processing is performed on the key receipt nodes through the intelligent big model to obtain multimodal key receipt node features, including:

[0014] Based on the item delivery scene video and the automatic interactive dialogue set, the intelligent big model is used to determine the key sign-off video frame set, key sign-off node dialogue segment, and node graph-text time sequence mapping information corresponding to the key sign-off node;

[0015] Determining key conversation semantic features corresponding to the key sign-off node conversation segment and key conversation visual features corresponding to the key sign-off video frame set;

[0016] Based on the mapping of the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information to the same target semantic feature space, the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information are subjected to multimodal information alignment processing in the target semantic feature space to obtain multimodal key sign-off node features.

[0017] In a feasible implementation, the structured event conversion process is performed using the item receipt dialogue event template based on the multimodal key receipt node features to obtain the item receipt scene description event, including:

[0018] Based on the multimodal key sign-off nodes, an intelligent big model is used to determine multiple rounds of key sign-off node conversation summaries, key sign-off node video clips, key sign-off node details, and item sign-off event description information;

[0019] The key receipt node dialogue summary, key receipt node video clip, key receipt node details and the item receipt event description information are filled into the item receipt dialogue event template to obtain an item receipt scene description event including multiple key receipt nodes.

[0020] In a feasible implementation manner, generating an item receipt autonomous response record based on the item receipt scenario description event includes:

[0021] Based on the item receipt scenario description event, a smart big model is used to determine multiple rounds of key receipt node dialogues, and key video clip links or detailed descriptions of key nodes are embedded in the key receipt node dialogues;

[0022] An item receipt autonomous response record is generated based on the key receipt node dialogue and in accordance with the dialogue sequence of the key receipt node.

[0023] In a feasible implementation manner, after pushing the autonomous response record of the item receipt to the user terminal, the method includes:

[0024] In response to a request to view the autonomous response record of the item receipt, an intelligent response record interface is loaded based on the autonomous response record of the item receipt, and the intelligent response record interface includes a key receipt node playback area and a response record display area.

[0025] In a feasible embodiment, the method further includes:

[0026] In response to a dialogue selection operation for a reference dialogue item on the intelligent response record interface, determining a reference key receipt node and a reference key node detailed description corresponding to the reference dialogue item;

[0027] The reference key video segment corresponding to the reference key video segment link is loaded in the key receipt node playback area, and the detailed description of the key node is displayed in the response record display area.

[0028] In a second aspect, an embodiment of this specification provides an item receipt processing device, the device comprising:

[0029] The monitoring module is used to monitor the delivery scene video in the home scene, and to communicate with the delivery personnel through the intelligent big model in the delivery scene to check the delivery process;

[0030] An acquisition module, configured to acquire a set of automatic interactive dialogues corresponding to the item receipt processing process;

[0031] a recording module configured to generate structured events based on the automatic interactive dialogue set and the item delivery scene video using an intelligent large model to obtain an item receipt scene description event, and generate an item receipt autonomous response record based on the item receipt scene description event;

[0032] The recording module is used to push the autonomous response record of the item receipt to the user terminal.

[0033] In a feasible implementation, the generating of structured events based on the automatic interactive dialogue set and the item delivery scene video by the intelligent big model to obtain the item receipt scene description event includes:

[0034] Perform semantic recognition on the automatic interactive dialogue set through an intelligent big model to determine multiple key sign-off nodes;

[0035] Based on the item delivery scene video and the automatic interactive dialogue set, the key sign-off nodes are mapped and aligned with multimodal node information using an intelligent big model to obtain multimodal key sign-off node features;

[0036] Based on the multimodal key receipt node features, an item receipt dialogue event template is used to perform structured event conversion processing to obtain an item receipt scene description event.

[0037] In a feasible implementation, based on the item delivery scene video and the automatic interactive dialogue set, the multimodal node information mapping and alignment processing is performed on the key receipt nodes through the intelligent big model to obtain multimodal key receipt node features, including:

[0038] Based on the item delivery scene video and the automatic interactive dialogue set, the intelligent big model is used to determine the key sign-off video frame set, key sign-off node dialogue segment, and node graph-text time sequence mapping information corresponding to the key sign-off node;

[0039] Determining key conversation semantic features corresponding to the key sign-off node conversation segment and key conversation visual features corresponding to the key sign-off video frame set;

[0040] Based on the mapping of the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information to the same target semantic feature space, the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information are subjected to multimodal information alignment processing in the target semantic feature space to obtain multimodal key sign-off node features.

[0041] In a feasible implementation, the structured event conversion process is performed using the item receipt dialogue event template based on the multimodal key receipt node features to obtain the item receipt scene description event, including:

[0042] Based on the multimodal key sign-off nodes, an intelligent big model is used to determine multiple rounds of key sign-off node conversation summaries, key sign-off node video clips, key sign-off node details, and item sign-off event description information;

[0043] The key receipt node dialogue summary, key receipt node video clip, key receipt node details and the item receipt event description information are filled into the item receipt dialogue event template to obtain an item receipt scene description event including multiple key receipt nodes.

[0044] In a feasible implementation manner, generating an item receipt autonomous response record based on the item receipt scenario description event includes:

[0045] Based on the item receipt scenario description event, a smart big model is used to determine multiple rounds of key receipt node dialogues, and key video clip links or detailed descriptions of key nodes are embedded in the key receipt node dialogues;

[0046] An item receipt autonomous response record is generated based on the key receipt node dialogue and in accordance with the dialogue sequence of the key receipt node.

[0047] In a feasible implementation manner, after pushing the autonomous response record of the item receipt to the user terminal, the method includes:

[0048] In response to a request to view the autonomous response record of the item receipt, an intelligent response record interface is loaded based on the autonomous response record of the item receipt, and the intelligent response record interface includes a key receipt node playback area and a response record display area.

[0049] In a feasible embodiment, the device further includes:

[0050] In response to a dialogue selection operation for a reference dialogue item on the intelligent response record interface, determining a reference key receipt node and a reference key node detailed description corresponding to the reference dialogue item;

[0051] The reference key video segment corresponding to the reference key video segment link is loaded in the key receipt node playback area, and the detailed description of the key node is displayed in the response record display area.

[0052] In a third aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0053] In a fourth aspect, an embodiment of this specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0054] The beneficial effects of the technical solutions provided by some embodiments of this specification include at least:

[0055] In one or more embodiments of this specification, by monitoring delivery scene videos in real time within a home environment, combining automated dialogue interactions between an intelligent large-scale model and delivery personnel, and collecting and forming a collection of automated interactive dialogues, a comprehensive and automated recording of the item receipt process is achieved through multimodal data fusion and structured event generation. This not only significantly improves the efficiency and accuracy of the receipt process, but also greatly simplifies the user's review and verification of the receipt process details through multimodal key node mapping and structured data display, allowing users to intuitively and conveniently obtain key information, thereby enhancing the security and user experience of the entire smart home system. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0057] Figure 1 This is a flowchart of a method for processing item receipt provided in an embodiment of this specification;

[0058] Figure 2 This is a schematic diagram of an intelligent response recording interface provided by an embodiment of this specification;

[0059] Figure 3 This is a schematic diagram of a flow chart of structured event generation provided by an embodiment of this specification;

[0060] Figure 4 This is a flowchart of a multi-modal key sign-off node feature processing provided by an embodiment of this specification;

[0061] Figure 5 This is a flow chart of a structured event conversion process provided by an embodiment of this specification;

[0062] Figure 6 This is a flow chart of generating an autonomous response record for item receipt provided by an embodiment of this specification;

[0063] Figure 7 This is a schematic diagram of the structure of an item receipt processing device provided in an embodiment of this specification;

[0064] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification;

[0065] Figure 9 This is a schematic diagram of the structure of the operating system and user space provided in the embodiments of this specification;

[0066] Figure 10 yes Figure 9 The architecture diagram of the Android operating system;

[0067] Figure 11 yes Figure 9 Architecture diagram of the IOS operating system. DETAILED DESCRIPTION

[0068] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this specification.

[0069] In the description of this specification, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units inherent to these processes, methods, products or devices. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to the specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship.

[0070] In recent years, with the rapid development of smart home and Internet of Things technologies, traditional methods of handling item receipts are faced with problems such as low manual operation efficiency and untimely information feedback. This manual proposes a method for handling item receipts. By introducing automated and intelligent solutions, it gradually explores the use of artificial intelligence, big data, and multimodal information fusion processing to achieve real-time monitoring of delivery scenarios and timely feedback, interaction, and data integration to users, thereby improving the accuracy of the entire receipt process and user experience.

[0071] The present specification is described in detail below with reference to specific embodiments.

[0072] In one embodiment, Figure 1As shown, a method for processing item receipt is proposed. This method can be implemented using a computer program and run on an item receipt processing device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone tool application. The item receipt processing device can be an electronic device, including but not limited to: a smart doorbell, a smart door lock, a smart camera, a personal computer, a tablet computer, a handheld device, an in-vehicle device, a wearable device, a computing device, or other processing device connected to a wireless modem. Terminal devices can be called different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, electronic device in 5G network or future evolution network, etc.

[0073] Specifically, the method for handling the receipt of the item includes:

[0074] S102: Monitoring the item delivery scene video in the home scene, and conducting the item receipt process with the item delivery personnel through the intelligent large model in the item delivery scene;

[0075] Home scene monitoring: refers to the use of smart home devices (such as smart cameras, doorbells, sensors, etc.) to monitor the environment such as the front door of the home in real time and capture image data related to the delivery of items.

[0076] Item delivery scene video: refers to video data recorded during the package delivery process, typically covering the delivery personnel entering the home, interacting with residents or equipment, and placing or signing for the package.

[0077] Intelligent large model: obtained by adapting the basic large language model (LLM) to the item receipt processing scenario.

[0078] Item receipt processing: refers to the actions and processes of completing package confirmation, rejection or exception handling during the interaction with the delivery personnel.

[0079] Schematically, smart home devices continuously monitor the home environment. When they detect a person approaching or a package being delivered, they automatically start recording and capturing relevant item delivery scene videos. The system uses a video analysis algorithm to determine that the current scene is a "delivery scene." Once confirmed, the system uses a built-in intelligent large model to initiate an automatic dialogue with the delivery person, such as through voice interaction or text prompts on the device. During the conversation, the system will proactively inquire about the delivery status of the package, such as "Is the package properly delivered?" and continue the subsequent process based on the delivery person's answer (voice or text).

[0080] For example: When a courier arrives at a user's door, the smart camera captures their image, and the system detects "delivery behavior" and automatically initiates a voice conversation: "Hello, has your package been delivered?" After the courier responds, the system records the interaction and uses it as a data source for subsequent processing.

[0081] S104: Obtaining an automatic interactive dialogue set corresponding to the item receipt processing process;

[0082] Automatic Interaction Conversation Collection: This refers to the collection of all automatically generated conversations between the intelligent big model and the delivery personnel during the item receipt process. This collection includes not only the conversations of both parties, but also the corresponding timestamps, speech-to-text transcripts, and metadata related to the conversation.

[0083] In this specification, considering that users will later view the item receipt video to understand the entire item receipt process, there is a tedious problem in the process of users directly searching and replaying long videos. This solution introduces an intelligent large model to perform in-depth processing of item receipt. Specifically, after the intelligent large model initially generates the item receipt scene description event, it will perform deep semantic analysis and feature fusion on the multimodal data (including the automatic interactive dialogue collection and the delivery scene video), thereby extracting the core information and video highlights of the key receipt nodes to generate the item receipt scene description event. Next, the system uses this extracted key information to generate a refined autonomous response record. The record displays the summary, details and corresponding video prompts of the key receipt nodes in the form of a dialogue, so that users can quickly understand the receipt process without manually searching the entire video content. In some embodiments, if the user needs more detailed on-site information, they can jump directly to the corresponding video clip by clicking on the video prompt in the record, realizing the linked playback of text and video. In this way, the intelligent large model enables the autonomous item receipt scene to provide an autonomous response record for item receipt, which not only greatly reduces the burden of users watching long videos, but also improves the intuitiveness and interactive convenience of the overall information presentation.

[0084] S106: Using the intelligent big model, a structured event is generated based on the automatic interactive dialogue set and the item delivery scene video to obtain an item receipt scene description event, and an item receipt autonomous response record is generated based on the item receipt scene description event;

[0085] The intelligent large model is obtained by adapting the task scenario to the item receipt scenario based on the trained basic large model (LLM);

[0086] Optional basic large language models include but are not limited to: GPT series large models, Tongyi Qianwen large model, DeepSeek large model, Wenxin Yiyan large model, etc.

[0087] Structured event generation: refers to the use of intelligent big models to extract, fuse and map multimodal data (conversation collections and videos) to generate standardized and structured data records. This record describes in detail the situation at each key node in the item receipt process.

[0088] Item receipt scenario description event: refers to a structured record containing information such as key receipt node conversation summaries, corresponding video clips, node details, and overall event description.

[0089] Item receipt autonomous response record: refers to the record generated by the system based on scene description events and presented in the form of a dialogue. The record reflects the interaction content between the intelligent big model and the delivery personnel at the key receipt nodes.

[0090] Schematically, from the collection of automatic interactive dialogues, an intelligent big model is used to identify and extract key text information such as "confirm receipt," "rejection," and "exception handling." The video of the item delivery scene is framed and image recognized to extract key frame video features related to receipt (such as the delivery personnel's actions when signing, the status of the package, etc.). Contextual information is synthesized based on the key text information and key frame video features. Then, a preset item receipt processing event template is called, which includes multiple fields, such as a summary of the key receipt node dialogue, a video clip link, a node detail description, and an overall event description. The intelligent big model is then used to extract event element fields from the contextual information, and the event element fields are filled into the preset item receipt processing event template to generate an item receipt scene description event. The structured data of the item receipt scene description event can accurately describe each key link in the entire receipt process.

[0091] Then, based on the generated item receipt scenario description events, the intelligent big model converts key conversation information into a conversation record between the agent and the delivery person according to a preset dialogue template. The conversation record presents the interactive content of key receipt nodes. Optionally, a video clip link or preview prompt can be embedded in each round of character dialogue in the conversation record. Users can click to directly view the corresponding video clip, achieving a text-video linkage effect.

[0092] Example: Assume that during a delivery process, the system identifies the key node as "Confirmation of Receipt." After multimodal alignment, the scenario description events may include:

[0093] Conversation excerpt: "The delivery person confirmed the package had been delivered and placed it at the door."

[0094] Video clip: "Keyframes 10:05-10:07 show the package placement action."

[0095] Node details: "The package packaging is intact and the lighting at the signing site is good."

[0096] Event description: "The system detected that the courier entered the signing area at 10:05 and then completed the package delivery confirmation."

[0097] Based on this event, the generated autonomous response record is as follows:

[0098] [Agent]: Hello, the system detected that your package was confirmed to have been received at 10:05. (Click to view the confirmation video clip)

[0099] [Delivery staff]: The package has been safely placed at your door, please check it.

[0100] S108: Pushing the autonomous response record of the item receipt to the user terminal.

[0101] User terminal: refers to the terminal device on which the user receives notifications and records, including but not limited to smartphones, tablets or PCs.

[0102] Push: It can be understood that the system transmits the generated item receipt autonomous response record to the user terminal in a timely manner through network communication, ensuring that the user can view the item receipt processing results in real time.

[0103] Indicatively, the generated autonomous response record for item receipt (including text conversations, video preview prompts, and related metadata) is encapsulated in a standardized format to obtain response record data, ensuring that the data adapts to the display requirements of the target user-side application. The response record data is transmitted to the user side through a preset network interface or push service (for example, using an HTTP API, a message queue, or a dedicated push service). After the user terminal receives the push record, the autonomous response record is displayed through the corresponding application interface. The user can directly click on the video link or preview icon in the record to achieve real-time playback of video clips of key receipt nodes and further verify the details of the receipt process.

[0104] In a feasible implementation manner, after pushing the autonomous response record of the item receipt to the user terminal, the method includes:

[0105] In response to the user's request to view the autonomous response record of the item receipt, an intelligent response record interface is loaded based on the autonomous response record of the item receipt. The intelligent response record interface includes a key receipt node playback area and a response record display area.

[0106] Indicative, such as Figure 2 As shown, Figure 2This is a diagram of the Smart Response Recording interface, which illustrates the corresponding Smart Response Recording interface for smart home devices. The upper half of the Smart Response Recording interface is the playback area for key receipt nodes. Typically, this area records video footage of corresponding key nodes, showing the scene outside the door or the presence of the delivery person during the item receipt process. The lower half is the "Smart Response Recording" dialog box, which displays the key conversation content between the smart doorbell (through its associated large model agent) and the delivery person in the form of chat bubbles.

[0107] From the dialogue bubble, you can see the following information about the key items received:

[0108] The agent asks: "Excuse me, are you XX Express?"

[0109] The delivery person replied: "Yes, I'm here to pick it up."

[0110] The intelligent agent provides information such as the express delivery location and pickup code;

[0111] Finally, the delivery person responded that the package had been picked up, and the agent politely ended the conversation.

[0112] This part of the conversation during the key item receipt process is designed to allow users to quickly understand the key exchanges between them and the delivery personnel, and to grasp the details of the receipt without having to replay the entire video.

[0113] In a feasible implementation manner, after the user terminal records the intelligent response record interface, the method further includes:

[0114] A2: In response to a user selecting a dialogue item on the smart answer record interface, determining a reference key receipt node and a reference key node detailed description corresponding to the reference dialogue item;

[0115] For example, when a user is viewing a "smart response record," they may need to learn more about the key video clip corresponding to a certain conversation node. To this end, the system provides clickable or selectable interactive portals for the key conversation items in the conversation record. The interactive portals may integrate references to key sign-off nodes and detailed descriptions of the key nodes.

[0116] Reference conversation items: These can be understood as key conversation items presented in the conversation record. These are conversations marked as being significantly relevant to the item receipt process or potentially of interest to the user. Examples include key phrases like "Please confirm if the package is damaged?" and "Will you sign for the package?" Each reference conversation item is associated with one or more key receipt nodes, along with detailed descriptions and video clips, when the response record is generated.

[0117] Conversation Selection: While viewing a Smart Answer record, users may be interested in a particular conversation and want to learn more about its context or replay the corresponding video scene. Users can select the "reference conversation item" by clicking, touching, or hovering their mouse. This action triggers the system to retrieve detailed information about the corresponding node.

[0118] In an illustrative manner, according to the dialogue item selected by the user, the reference key sign-off node of the dialogue item stored in the backend data is queried to determine the reference key node detailed description corresponding to the reference key sign-off node;

[0119] A4: Load the reference key video clip corresponding to the reference key video clip link in the key receipt node playback area, and display the detailed description of the key node in the response record display area.

[0120] Key Signoff Node Playback Area: The interface typically features a "Video Playback" or "Key Node Playback" area (possibly above, next to, or in a pop-up window). Once the system identifies a key signoff node for reference, it can use its timestamp or video clip link to locate and play the video in this area.

[0121] Each key sign-off node is then associated with one or more video time segments, such as "Start Time: 00:01:20, End Time: 00:01:45." The system then jumps the video player's progress bar or cursor to that time range and begins playing or preloading that segment. This allows users to simply click on a reference dialogue item to directly view the corresponding actual video scene, eliminating the need to search through the entire video.

[0122] Furthermore, detailed descriptions of key nodes are displayed in the response record display area. Synchronously with video playback, the system displays detailed descriptions of previously determined key nodes in the response record display area or in an adjacent description panel. This description can appear in the form of text, images, labels, etc., to help users intuitively understand the key points of the scene corresponding to the current video clip.

[0123] In the examples of this specification, by monitoring delivery scenes in real time within a home environment, combining automated dialogue interactions between an intelligent large-scale model and delivery personnel, and collecting and generating a collection of automated interactive dialogues, a comprehensive and automated recording of the item receipt process is achieved through multimodal data fusion and structured event generation. This not only significantly improves the efficiency and accuracy of the receipt process, but also greatly simplifies the user's review and verification of the receipt process details through multimodal key node mapping and structured data display, allowing users to intuitively and conveniently obtain key information, thereby enhancing the overall security and user experience of the smart home system.

[0124] See Figure 3 , Figure 3 This is a flow chart of a structured event generation process proposed in this specification. Specifically, the process of generating a structured event based on the automatic interactive dialogue set and the item delivery scene video using the intelligent big model to obtain an item receipt scene description event may include:

[0125] S202: Performing semantic recognition on the automatic interactive dialogue set using an intelligent big model to determine multiple key sign-off nodes;

[0126] Key receipt nodes: These are the conversations or event nodes during the delivery process, representing important steps such as the delivery staff's door-to-door visit, receipt confirmation, rejection, and exception handling. Key receipt nodes are usually the nodes that users pay the most attention to during the item receipt process.

[0127] Schematically, the automated interactive dialogue collection is input into the intelligent big model. Based on its pre-trained semantic understanding capabilities, the intelligent big model performs semantic encoding and feature extraction on the text content. The intelligent big model captures key semantic signals such as "Express delivery, please sign for it," "Confirmed receipt," "Package delivered," "Rejected," and "Package abnormality." Based on these key semantic signals, multiple key sign-off nodes are determined.

[0128] For example, the automated interaction transcript contains the following two conversations:

[0129] "[Agent]: Has the package been delivered safely?" (Time: 10:05)

[0130] "[Delivery Personnel]: Yes, the package has been left at the door." (Time: 10:06)

[0131] After semantic recognition, the system will mark the second conversation as the "Confirm Receipt" key node and record the key information of the node (such as the node type is "Confirm Receipt", the timestamp is 10:06, and the summary is "The package has been placed in front of the door").

[0132] S204: Based on the item delivery scene video and the automatic interactive dialogue set, a multimodal node information mapping and alignment process is performed on the key receipt nodes using an intelligent big model to obtain multimodal key receipt node features;

[0133] Item delivery scene video: refers to the delivery process video recorded by smart home devices, which contains various information such as the actual operations of the delivery personnel, the status of the package, and the on-site environment.

[0134] Multimodal node information mapping alignment: Multimodal data fusion technology is used to temporally and semantically match key information in text (conversation) with visual information in video (such as key frames, action recognition results), thereby forming a unified multimodal feature describing the same event.

[0135] Multimodal key sign-off node features: refers to the comprehensive data structure in which each key sign-off node, after fusion processing, has both text semantic features and corresponding video visual features, which is used to comprehensively describe the event represented by the node.

[0136] Schematically, the intelligent large-scale model processes the video of the item delivery scene by framing it, extracting the visual features of each frame and recording the timestamp information of each frame to ensure that the video data and the conversation data are temporally aligned. Using the timestamps of the key sign-off nodes recorded in S202, the key sign-off nodes in the text are matched with the corresponding time periods in the video to achieve temporal alignment. Then, based on the attention mechanism, the semantic vector of the text and the visual features of the video frames are jointly encoded. The intelligent large-scale model automatically learns the correspondence between the text and the video, generating a multimodal key sign-off node feature. The multimodal key sign-off node feature includes: node type, text summary, visual feature vector, video key frame information, time range, etc., forming a structured data record.

[0137] In one possible implementation, please refer to Figure 4 , Figure 4 This is a flowchart of multimodal key sign-off node feature processing. The intelligent big model is used to specifically execute the multimodal node information mapping and alignment processing based on the item delivery scene video and the automatic interactive dialogue set. The multimodal key sign-off node features can be obtained by referring to the following method:

[0138] S302: Based on the item delivery scene video and the automatic interactive dialogue set, determine the key sign-off video frame set, key sign-off node dialogue segment, and node graph-text time sequence mapping information corresponding to the key sign-off node through the intelligent big model;

[0139] Item delivery scene video: refers to the delivery process video collected by smart home devices, which records information such as delivery personnel operations, package status, and on-site environment.

[0140] Automatic interactive dialogue collection: refers to all conversations between the intelligent big model and delivery personnel automatically recorded by the system, including text converted from speech, timestamps, and role identification.

[0141] Key signing nodes: nodes that are representative and have decision-making significance in the signing process, such as "package confirmation", "rejection", "exception handling", etc.

[0142] Key receipt video frame set: a set of image frames extracted from the delivery video that can reflect the key receipt node scenes.

[0143] Key sign-off node dialogue fragment: a dialogue content fragment directly related to the key sign-off node, captured from the automatic interaction dialogue set.

[0144] Node graph-text temporal mapping information: data that records the temporal correspondence between images (video frames) and text (dialogues), indicating which segment of dialogue corresponds to which time period in the video.

[0145] Schematically, the delivery scene video is framed and processed through the intelligent big model, each frame is extracted, and key frames closely related to the signing operation are selected using technologies such as target detection and action recognition to obtain a set of key signing video frames (such as frames when the delivery person places the package or confirms the signing); the intelligent big model is used to analyze the collection of automatic interactive dialogues, and semantic recognition technology is used to extract dialogue segments related to the signing node (such as sentences such as "Please confirm that the package has been placed" or "The package has been delivered") and retain the timestamps corresponding to the dialogue segments of the key signing node; the intelligent big model uses the time information recorded in the dialogue and video to match the key dialogue segments with the key frames of the corresponding time period in the video, and generate image-text temporal mapping information. This ensures that the text and visual content are accurately aligned in time.

[0146] S304: Determine key conversation semantic features corresponding to the key sign-off node conversation segment and key conversation visual features corresponding to the key sign-off video frame set;

[0147] Key conversation semantic features: Semantic representations (usually vectors or embeddings) extracted from conversation fragments at key sign-off nodes using natural language processing technology can capture the semantic intent and emotional information in the conversation.

[0148] Key dialogue visual features: Visual features extracted from a set of key sign-off video frames, reflecting the semantics and visual details of the image content.

[0149] Illustratively, for the key sign-off node conversation segments selected in S302, the text is encoded to obtain a set of high-dimensional vector representations as key conversation semantic features. Key conversation semantic features can reflect semantic information such as "confirmation of receipt" and "exception handling" in the conversation. For the set of key sign-off video frames extracted in S302, feature extraction is performed on each frame to obtain key conversation visual features.

[0150] S306: Based on the mapping of the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information to the same target semantic feature space, the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information are subjected to multimodal information alignment processing in the target semantic feature space to obtain multimodal key sign-off node features.

[0151] Target semantic feature space: a unified feature representation space in which features from different modalities (text, image) can be directly compared and fused, usually through projection or mapping layers.

[0152] Multimodal information alignment processing: Match and integrate features from text and vision according to the common semantic context, eliminate inconsistencies between modalities, and enable them to be effectively integrated in the same semantic space.

[0153] Schematically, the intelligent large model can use a fully connected layer to map the key conversation semantic features and key conversation visual features extracted in S304 to the same target semantic feature space. This process can be achieved by training a network with shared parameters so that the features of the two modalities are comparable in the same dimension. In the target semantic feature space, the mapped text and visual features are aligned in combination with the node graph and text temporal mapping information obtained in S302. For example, the attention mechanism is used to perform weighted fusion of text and image information to ensure that the two features are highly consistent when representing the same key node. Finally, the fused features are integrated into a unified vector or data structure to obtain multimodal key sign-off node features. The data structure also contains text semantics, visual details, and temporal mapping information to fully describe the multimodal information of the key sign-off node.

[0154] In this manual, through a multi-step, multi-modal information alignment and fusion method, the system is able to extract and integrate key information from text and video during the signing process, providing detailed and accurate data support for subsequent structured event generation and autonomous response records, while greatly improving the user's understanding of the details of the signing process and verification efficiency.

[0155] S206: Based on the multimodal key receipt node features, an item receipt dialogue event template is used to perform structured event conversion processing to obtain an item receipt scene description event.

[0156] Item Receipt Dialogue Event Template: A predefined structured data format used to store information about key nodes in the item receipt process. The Item Receipt Dialogue Event Template may contain fields such as a conversation summary, a video clip link, detailed node descriptions, and an overall event description.

[0157] Structured event conversion processing: Map and fill the data extracted from multimodal features according to a predetermined template format to generate standardized event records for easy storage, retrieval, and display.

[0158] Indicatively, a template for item receipt event is pre-designed, which divides the information of all key nodes into multiple parts, such as:

[0159] Conversation summary: briefly describe the core conversation content at key points;

[0160] Video clip link: a link to a clip or preview image corresponding to a key moment in the video;

[0161] Node Detailed Description: Detailed explanation of the background information of key nodes, detected anomalies and other supplementary explanations;

[0162] Overall description of the event: a summary of the entire signing process.

[0163] The multimodal key sign-off node features generated in S204 are then mapped according to the requirements of each field. For example, a text summary is entered into the "Conversation Summary" field, key frames and time information extracted from the video are converted into "Video Clip Links," and detailed visual and semantic descriptions of the node are entered into the "Node Detailed Description" field. After mapping and data filling, a complete item sign-off scenario description event is generated. This event is recorded as structured data and can fully describe the key sign-off process from automatic interaction to video playback.

[0164] In the embodiments of this specification, through the semantic recognition of the intelligent big model, multiple key signing nodes are extracted and determined from the automatic interactive dialogue, and key information is extracted for the entire item signing process. Then, the time alignment and multimodal fusion of the video data and the dialogue data are used to map and align the text and visual information of each key node to generate a multimodal feature data containing rich information. Finally, through the preset event template, the multimodal features are structured and converted to generate a standardized item signing scene description event. This event comprehensively and intuitively records the key links of the signing process, providing a solid data foundation for subsequent automatic response records and user interface displays. This solution makes full use of the advantages of the intelligent big model in natural language processing and multimodal information fusion, and realizes the efficient conversion from original dialogue and video data to structured event records, which not only improves the accuracy and real-time nature of the information, but also significantly optimizes the user's backtracking and verification experience of the item signing process.

[0165] In one possible implementation, please refer to Figure 5 , Figure 5 This is a flow chart of structured event conversion. Specifically, the structured event conversion process is performed based on the multimodal key receipt node features using the item receipt dialogue event template to obtain the item receipt scene description event. The following method can be used for reference:

[0166] S402: Determine, based on the multimodal key sign-off nodes, a multi-round key sign-off node dialogue summary, a key sign-off node video clip, key sign-off node details, and item sign-off event description information using an intelligent big model;

[0167] A multi-round dialogue summary for a key delivery node: This is a comprehensive, concise text summary of the multiple dialogue rounds involved in the same key node, capturing the core interaction content of that node. For example, the node "Confirming package delivery" may involve multiple confirmation conversations, and the summary should include all key information.

[0168] Key delivery node video clips: Continuous video segments related to key delivery nodes extracted from the video. This video clip accurately reflects the scene at that node, such as the specific operation process of the delivery staff placing the package in front of the door.

[0169] Details of key signing nodes: For each key signing node, in addition to the conversation summary and video clips, detailed information is also included, such as a description of the environment at the time, detected anomalies (such as package damage, delivery anomalies, etc.), time information, and other supplementary annotations.

[0170] Item receipt event description information: It is an overall description of the entire item receipt process, usually based on the summary of all key nodes, providing a comprehensive and general event summary.

[0171] Schematically, the obtained multimodal key sign-off node features are used as input to the intelligent big model. Based on pre-set tasks or prompts, the model automatically analyzes the multi-round conversation data for each key node, extracts the key information from the conversation, and generates a refined summary to obtain the multi-round key sign-off node conversation summary. Simultaneously, the model identifies and extracts the corresponding key sign-off node video clips from the multimodal features, determining which segment of the video reflects the core content of the event at that node. The intelligent big model further combines the semantic and visual details of the node to generate key sign-off node details for that node. These key sign-off node details may include environmental descriptions, anomaly detection, and operational details. Finally, based on the multi-round key sign-off node conversation summaries, key sign-off node video clips, and key sign-off node details, the intelligent big model automatically generates a comprehensive description of the item sign-off event as the item sign-off event description information, which provides key information for the entire delivery process.

[0172] S404: Fill the key receipt node dialogue summary, key receipt node video clip, key receipt node details and the item receipt event description information into the item receipt dialogue event template to obtain an item receipt scene description event including multiple key receipt nodes.

[0173] The item receipt dialogue event template is a predefined standardized template that aims to uniformly describe the key nodes in the item receipt process. The template contains multiple fields, such as:

[0174] Conversation summary field: stores the conversation summary of each key node.

[0175] Video clip field: stores the corresponding key video clip or its link.

[0176] Node details field: stores detailed description information of each node.

[0177] Event description field: overall description of the content of the entire receipt event.

[0178] Structured event conversion processing: refers to the data mapping and filling of unstructured or semi-structured multimodal information according to preset templates, and ultimately generates standardized event records that are easy to store and retrieve.

[0179] Schematically, the multi-round key receipt node conversation summaries, video clips, node details and overall event description information generated in S402 are data mapped. For each key node, its summary, video clips and details are filled into the corresponding fields in the template in turn. At the same time, the information of all key nodes is summarized and filled into the overall event description field to form a complete and structured item receipt scene description event.

[0180] In this manual, an intelligent large-scale model extracts multi-round conversation summaries, video clips, detailed descriptions, and event descriptions from multimodal key receipt node features, enabling deep mining and refinement of key node information. Furthermore, by populating this information into predefined item receipt conversation event templates, a structured, standardized item receipt scenario description event is generated, covering multiple key receipt nodes. This provides a comprehensive and intuitive record for subsequent user presentation, data storage, and problem tracing. The entire process achieves in-depth processing and standardized conversion of multimodal data, not only enabling efficient information integration but also significantly improving users' understanding and verification efficiency of key nodes in the delivery and receipt process.

[0181] For further information, see Figure 6 , Figure 6 This is a flow chart of generating an autonomous response record for item receipt. Specifically, the following methods can be used to generate an autonomous response record for item receipt based on the item receipt scenario description event:

[0182] S502: Based on the item receipt scenario description event, a smart big model is used to determine multiple rounds of key receipt node dialogues, and key video clip links and detailed descriptions of key nodes are embedded in the key receipt node dialogues;

[0183] Multiple rounds of dialogues at key sign-off nodes: This refers to the dialogues and exchanges involving key sign-off nodes at different time points during the item sign-off process. These dialogues cover core links such as confirmation, exception handling, and rejection.

[0184] Key video clip link: It is a video clip identifier extracted from the scene video and associated with a specific sign-off node. It is usually presented as a link or preview image, allowing users to directly play back the corresponding video content after clicking it.

[0185] Detailed description of key nodes: It is a text description of the background, environment, abnormal conditions, etc. of each key receipt node, aiming to provide users with more comprehensive information supplement.

[0186] Schematically, the previously generated item receipt scenario description event is used as input. This event contains multimodal data (conversation summaries, video clips, detailed descriptions, etc.) for multiple key nodes. Using an intelligent big model, the input event is semantically parsed, identifying the individual segments that constitute multiple rounds of key receipt conversations. Based on timestamps and semantic clues, the model reconstructs the complete conversation flow, such as the continuous interaction from "Please confirm that the package has been delivered?" to "The package has been placed at the door." For each key round of dialogue, the model automatically embeds a link to the key video clip corresponding to that node in the conversation text, such as inserting a "(Click to view video)" prompt at the corresponding point in the conversation. Furthermore, detailed descriptions of key nodes (such as "The lighting on site is good, and the package appears intact") are appended to the conversation description or displayed as notes. After processing, each key receipt conversation not only contains the original text exchange, but also integrates a video playback link and detailed descriptions, forming a multimodal conversation entry that is easily integrated into the final response record.

[0187] S504: Generate an item receipt autonomous response record based on the key receipt node dialogue and the dialogue sequence of the key receipt node.

[0188] Key receipt node dialogue: The multi-round dialogue entries containing video links and detailed descriptions obtained after processing in S502 reflect the communication content at each key node in the item receipt process.

[0189] Conversation sequence: refers to the order in which these conversation items are arranged according to the actual time of occurrence during the sign-off process, reflecting the time process and logical order of events.

[0190] Item receipt autonomous response record: The final record generated shows the key information exchange between the intelligent big model and the delivery personnel in the form of a dialogue. Users can directly view, interact, and replay the corresponding video, thereby realizing convenient backtracking and verification of the entire receipt process.

[0191] For example, all conversation items are sorted chronologically based on the timestamp information in each key sign-off conversation, ensuring that the final record matches the actual process sequence. The sorted conversation items are then combined into a complete autonomous response record for the item's sign-off according to a pre-set conversation format. This record is typically presented as a conversation thread with clear role distinctions and time sequence markers.

[0192] In this manual, based on the description events of the item receipt scenario and the multimodal key receipt node features, an intelligent large model is used to conduct in-depth semantic analysis of the conversations at each key node, and corresponding key video clip links and detailed descriptions are embedded to make the communication content of each node richer and more intuitive. According to the chronological order of the conversations, all key node conversation items are integrated and sorted to generate the final item receipt autonomous response record. This record is presented in the form of a conversation, which not only retains the original interaction content, but also realizes the linkage display of text and video by embedding multimodal information, significantly improving the efficiency and accuracy of users viewing and verifying the item receipt process.

[0193] The following will be combined Figure 7 , the article receipt processing device provided in the embodiment of this specification is introduced in detail. It should be noted that, Figure 7 The article receipt processing device shown is used to execute this instruction Figures 1 to 6 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to this specification. Figures 1 to 6 The embodiment shown.

[0194] See Figure 7 , which shows a schematic diagram of the structure of the item receipt processing device according to an embodiment of this specification. The item receipt processing device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the item receipt processing device 1 includes a monitoring module 11, an acquisition module 12, and a recording module 13, which are specifically used to:

[0195] Monitoring module 11 is used to monitor the item delivery scene video in the home scene, and conduct the item receipt processing process with the item delivery personnel through the intelligent large model in the item delivery scene;

[0196] An acquisition module 12 is configured to acquire a set of automatic interactive dialogues corresponding to the item receipt processing process;

[0197] A recording module 13 is configured to generate a structured event based on the automatic interactive dialogue set and the item delivery scene video using an intelligent big model to obtain an item receipt scene description event, and generate an item receipt autonomous response record based on the item receipt scene description event;

[0198] The recording module is used to push the autonomous response record of the item receipt to the user terminal.

[0199] In a feasible implementation, the generating of structured events based on the automatic interactive dialogue set and the item delivery scene video by the intelligent big model to obtain the item receipt scene description event includes:

[0200] Perform semantic recognition on the automatic interactive dialogue set through an intelligent big model to determine multiple key sign-off nodes;

[0201] Based on the item delivery scene video and the automatic interactive dialogue set, the key sign-off nodes are mapped and aligned with multimodal node information using an intelligent big model to obtain multimodal key sign-off node features;

[0202] Based on the multimodal key receipt node features, an item receipt dialogue event template is used to perform structured event conversion processing to obtain an item receipt scene description event.

[0203] In a feasible implementation, based on the item delivery scene video and the automatic interactive dialogue set, the multimodal node information mapping and alignment processing is performed on the key receipt nodes through the intelligent big model to obtain multimodal key receipt node features, including:

[0204] Based on the item delivery scene video and the automatic interactive dialogue set, the intelligent big model is used to determine the key sign-off video frame set, key sign-off node dialogue segment, and node graph-text time sequence mapping information corresponding to the key sign-off node;

[0205] Determining key conversation semantic features corresponding to the key sign-off node conversation segment and key conversation visual features corresponding to the key sign-off video frame set;

[0206] Based on the mapping of the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information to the same target semantic feature space, the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information are subjected to multimodal information alignment processing in the target semantic feature space to obtain multimodal key sign-off node features.

[0207] In a feasible implementation, the structured event conversion process is performed using the item receipt dialogue event template based on the multimodal key receipt node features to obtain the item receipt scene description event, including:

[0208] Based on the multimodal key sign-off nodes, an intelligent big model is used to determine multiple rounds of key sign-off node conversation summaries, key sign-off node video clips, key sign-off node details, and item sign-off event description information;

[0209] The key receipt node dialogue summary, key receipt node video clip, key receipt node details and the item receipt event description information are filled into the item receipt dialogue event template to obtain an item receipt scene description event including multiple key receipt nodes.

[0210] In a feasible implementation manner, generating an item receipt autonomous response record based on the item receipt scenario description event includes:

[0211] Based on the item receipt scenario description event, a smart big model is used to determine multiple rounds of key receipt node dialogues, and key video clip links or detailed descriptions of key nodes are embedded in the key receipt node dialogues;

[0212] An item receipt autonomous response record is generated based on the key receipt node dialogue and in accordance with the dialogue sequence of the key receipt node.

[0213] In a feasible implementation manner, after pushing the autonomous response record of the item receipt to the user terminal, the method includes:

[0214] In response to a request to view the autonomous response record of the item receipt, an intelligent response record interface is loaded based on the autonomous response record of the item receipt, and the intelligent response record interface includes a key receipt node playback area and a response record display area.

[0215] In a feasible embodiment, the device further includes:

[0216] In response to a dialogue selection operation for a reference dialogue item on the intelligent response record interface, determining a reference key receipt node and a reference key node detailed description corresponding to the reference dialogue item;

[0217] The reference key video segment corresponding to the reference key video segment link is loaded in the key receipt node playback area, and the detailed description of the key node is displayed in the response record display area.

[0218] In a third aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the above-mentioned method steps.

[0219] In a fourth aspect, an embodiment of this specification provides an electronic device, which may include: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.

[0220] It should be noted that the item receipt processing device provided in the above embodiment only uses the division of the above functional modules as an example when executing the item receipt processing method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the item receipt processing device provided in the above embodiment and the item receipt processing method embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0221] The serial numbers of the embodiments in this specification are for description only and do not represent the advantages or disadvantages of the embodiments.

[0222] The embodiment of this specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor as described above. Figures 1 to 6 The specific implementation process of the item receipt processing method in the embodiment shown can be found in Figures 1 to 6 The detailed description of the illustrated embodiment will not be repeated here.

[0223] This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1 to 6 The specific implementation process of the item receipt processing method in the embodiment shown can be found in Figures 1 to 6 The detailed description of the illustrated embodiment will not be repeated here.

[0224] Please refer to Figure 8 , which shows a block diagram of the structure of an electronic device provided by an exemplary embodiment of this specification. The electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.

[0225] The processor 110 may include one or more processing cores. The processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device and process data. Optionally, the processor 110 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communications chip.

[0226] The memory 120 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The operating system may be an Android system, including a system deeply developed based on the Android system, an IOS system developed by Apple, including a system deeply developed based on the IOS system or other systems. The data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat record data, etc.

[0227] See also Figure 9As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, and native and third-party applications run in the user space. In order to ensure that different third-party applications can achieve better operating results, the operating system allocates corresponding system resources to different third-party applications. However, the requirements for system resources in different application scenarios in the same third-party application are also different. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance. The operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.

[0228] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to open up data communication between third-party applications and the operating system so that the operating system can obtain the current scenario information of third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0229] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows: Figure 10As shown, the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380. The Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space. The Linux kernel layer 320 provides underlying drivers for various hardware components of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 340 provides major feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider management, package management, call management, resource management, and location management. The application layer 380 runs at least one application. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera applications, or third-party applications developed by third-party developers, such as games, instant messaging programs, and photo enhancement programs.

[0230] Taking the operating system as the IOS system as an example, the programs and data stored in the memory 120 are as follows: Figure 11As shown, the IOS system includes: a core operating system layer 420 (Core OS layer), a core service layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch Layer). The core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide functions closer to the hardware for use by the program framework located in the core service layer 440. The core service layer 440 provides system services and / or program frameworks required by applications, such as the foundation framework, account framework, advertising framework, data storage framework, network connection framework, geographic location framework, motion framework, etc. The media layer 460 provides applications with audio-visual interfaces, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and wireless playback (AirPlay) interfaces for audio and video transmission technologies. The touchable layer 480 provides various commonly used interface-related frameworks for application development. The touchable layer 480 is responsible for user touch interaction operations on electronic devices. For example, local notification service, remote push service, advertising framework, game tool framework, message user interface (UI) framework, user interface UIKit framework, map framework, etc.

[0231] exist Figure 11 Among the frameworks shown, those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480. The Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI. The classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.

[0232] Among them, the method and principle of implementing data communication between third-party applications and the operating system in the IOS system can be referred to the Android system, and this manual will not go into details here.

[0233] Among them, the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device. The output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the user using any suitable objects such as fingers and touch pens, and to display the user interface of each application. The touch screen display is usually provided on the front panel of the electronic device. The touch screen display can be designed as a full screen, a curved screen or a special-shaped screen. The touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in the embodiments of this specification.

[0234] In addition, those skilled in the art will understand that the structures of the electronic devices shown in the above figures do not limit the electronic devices. The electronic devices may include more or fewer components than shown, or may combine certain components, or arrange the components differently. For example, the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which are not described in detail here.

[0235] In the embodiments of this specification, the execution entity of each step can be the electronic device described above. Optionally, the execution entity of each step is the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems, and this embodiment of this specification does not limit this.

[0236] The electronic device of the embodiment of this specification may further be equipped with a display device, and the display device may be any device capable of realizing a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc. The user may use the display device on the electronic device to view displayed text, images, videos and other information. The electronic device may be a smart phone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio playback device, a video playback device, a notebook, a desktop computing device, a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, electronic clothing and the like.

[0237] exist Figure 8 In the electronic device shown, the processor 110 can be used to call the application stored in the memory 120 and specifically execute the item receipt processing method involved in one or more embodiments of the following description of this specification.

[0238] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.

[0239] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A method for processing item receipt, characterized in that: The method comprises: Monitor the delivery scene video in the home scene, and use the intelligent big model to communicate with the delivery personnel during the delivery process; Obtaining a set of automatic interactive dialogues corresponding to the item receipt processing process; Generate an item receipt scene description event by performing structured event generation based on the automatic interactive dialogue set and the item delivery scene video through an intelligent big model, and generate an item receipt autonomous response record based on the item receipt scene description event; Push the autonomous response record of the item receipt to the user end.

2. The method according to claim 1, wherein generating a structured event based on the automatic interactive dialogue set and the item delivery scene video using an intelligent big model to obtain an item receipt scene description event comprises: Perform semantic recognition on the automatic interactive dialogue set through an intelligent big model to determine multiple key sign-off nodes; Based on the item delivery scene video and the automatic interactive dialogue set, the key sign-off nodes are mapped and aligned with multimodal node information using an intelligent big model to obtain multimodal key sign-off node features; Based on the multimodal key receipt node features, an item receipt dialogue event template is used to perform structured event conversion processing to obtain an item receipt scene description event.

3. The method according to claim 2, wherein, based on the item delivery scene video and the automated interactive dialogue set, a multimodal node information mapping and alignment process is performed on the key sign-off nodes using an intelligent big model to obtain multimodal key sign-off node features, including: Based on the item delivery scene video and the automatic interactive dialogue set, the intelligent big model is used to determine the key sign-off video frame set, key sign-off node dialogue segment, and node graph-text time sequence mapping information corresponding to the key sign-off node; Determining key conversation semantic features corresponding to the key sign-off node conversation segment and key conversation visual features corresponding to the key sign-off video frame set; Based on the mapping of the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information to the same target semantic feature space, the key conversation semantic features, the key conversation visual features and the node graph-text timing mapping information are subjected to multimodal information alignment processing in the target semantic feature space to obtain multimodal key sign-off node features.

4. The method according to claim 2, wherein the structured event conversion process is performed using the item receipt dialogue event template based on the multimodal key receipt node features to obtain the item receipt scene description event, including: Based on the multimodal key sign-off nodes, an intelligent big model is used to determine multiple rounds of key sign-off node conversation summaries, key sign-off node video clips, key sign-off node details, and item sign-off event description information; The key receipt node dialogue summary, key receipt node video clip, key receipt node details and the item receipt event description information are filled into the item receipt dialogue event template to obtain an item receipt scene description event including multiple key receipt nodes.

5. The method according to claim 4, wherein generating an autonomous response record for item receipt based on the item receipt scenario description event comprises: Based on the item receipt scenario description event, a smart big model is used to determine multiple rounds of key receipt node dialogues, and key video clip links and detailed descriptions of key nodes are embedded in the key receipt node dialogues; An item receipt autonomous response record is generated based on the key receipt node dialogue and in accordance with the dialogue sequence of the key receipt node.

6. The method according to claim 1, after pushing the autonomous response record of the item receipt to the user terminal, further comprising: In response to a request to view the autonomous response record of the item receipt, an intelligent response record interface is loaded based on the autonomous response record of the item receipt, and the intelligent response record interface includes a key receipt node playback area and a response record display area.

7. The method according to claim 6, further comprising: In response to a dialogue selection operation for a reference dialogue item on the intelligent response record interface, determining a reference key receipt node and a reference key node detailed description corresponding to the reference dialogue item; The reference key video segment corresponding to the reference key video segment link is loaded in the key receipt node playback area, and the detailed description of the key node is displayed in the response record display area.

8. An article receipt processing device, characterized in that: Application and smart home device, the device includes: The monitoring module is used to monitor the delivery scene video in the home scene, and to communicate with the delivery personnel through the intelligent big model in the delivery scene to check the delivery process; An acquisition module, configured to acquire a set of automatic interactive dialogues corresponding to the item receipt processing process; a recording module configured to generate structured events based on the automatic interactive dialogue set and the item delivery scene video using an intelligent large model to obtain an item receipt scene description event, and generate an item receipt autonomous response record based on the item receipt scene description event; The recording module is used to push the autonomous response record of the item receipt to the user terminal.

9. A computer storage medium, characterized in that The computer storage medium stores a plurality of instructions, which are suitable for being loaded by a processor and executing the method steps according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the method steps according to any one of claims 1 to 7.