An information analysis method and device, electronic equipment and storage medium

By constructing a message database and using large models to automatically filter and analyze behavioral events of smart home devices, the problem of insufficient intelligence in existing technologies is solved, and efficient behavioral event filtering and analysis are achieved.

CN121117154BActive Publication Date: 2026-08-04HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU EZVIZ SOFTWARE CO LTD
Filing Date
2025-08-28
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing smart home entry devices typically only record a single event when they detect behavioral events at the door, requiring users to manually filter and analyze them to obtain answers, resulting in insufficient intelligence.

Method used

By building a message database, filtering candidate messages that meet the time range, rewriting the candidate messages to merge the identities of the same person, and using a large model to generate feedback results, automatic filtering and analysis are achieved.

Benefits of technology

It improves the intelligent screening and analysis capabilities of behavioral events, generates accurate feedback results, and reduces the need for manual operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121117154B_ABST
    Figure CN121117154B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an information analysis method and device, electronic equipment and storage medium, and relate to the technical field of smart home, the method comprising: in response to obtaining a target user question related to entry, determining a screening condition; from a message database, screening candidate messages that meet the screening condition; performing message rewriting on each candidate message obtained by screening to obtain each target message; constructing a prompt word containing at least the target user question, the target message, and the explanation content for the target message; inputting the prompt word into a predetermined large model, so that the large model generates a feedback result matched with the target user question in response to the prompt word. Through the present application, the behavior event can be automatically screened and analyzed, and the intelligent degree of screening and analysis of behavior events is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home technology, and in particular to an information analysis method, device, electronic device and storage medium. Background Technology

[0002] With the rapid development of smart home technology, smart entry devices have become an important part of modern home security and convenience. Types of smart entry devices include smart door locks, smart peepholes, and smart doorbells. These devices can identify and record behavioral events at the door, such as loitering, opening the door, and knocking.

[0003] Existing smart entry devices typically only record simple, single behavioral events when they detect them at the door. When users need answers to any questions related to entry, they often have to go through a large number of behavioral events one by one and manually filter and analyze them, resulting in insufficient intelligence. Summary of the Invention

[0004] The purpose of this application is to provide an information analysis method, apparatus, electronic device, and storage medium to automatically filter and analyze behavioral events, thereby improving the intelligence level of filtering and analyzing behavioral events. The specific technical solution is as follows:

[0005] In a first aspect, embodiments of this application provide an information analysis method, the method comprising:

[0006] In response to obtaining target user questions related to home visits, filtering criteria are determined; wherein, the filtering criteria include a time range obtained based on the intent represented by the target user questions;

[0007] From the message database, candidate messages that meet the filtering conditions are filtered; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door.

[0008] The selected candidate messages are rewritten to obtain target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person.

[0009] Construct prompt words that include at least the target user's question, each target message, and an explanation of each target message;

[0010] The prompt word is input into a predetermined large model, so that the large model responds to the prompt word and generates a feedback result that matches the target user's question.

[0011] Secondly, embodiments of this application provide an information analysis device, the device comprising:

[0012] The determination module is used to determine filtering conditions in response to obtaining target user questions related to home visits; wherein the filtering conditions include a time range obtained based on the intent represented by the target user questions;

[0013] The filtering module is used to filter candidate messages that meet the filtering conditions from the message database; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door.

[0014] The rewriting module is used to rewrite the candidate messages obtained from the screening to obtain the target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person.

[0015] A building module is used to build prompts that include at least the target user's question, the various target messages, and explanations for the various target messages;

[0016] An input module is used to input the prompt words into a predetermined large model, so that the large model responds to the prompt words and generates a feedback result that matches the target user's question.

[0017] Thirdly, embodiments of this application provide an electronic device, including: a memory for storing computer programs; and a processor for executing the program stored in the memory to implement any of the aforementioned information analysis methods.

[0018] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the aforementioned information analysis methods.

[0019] This application also provides a computer program product containing instructions that, when run on a computer, cause the computer to execute any of the information analysis methods described above.

[0020] Beneficial effects of the embodiments in this application:

[0021] This application provides an information analysis method that includes a message database. Each message in the message database corresponds to an image. Each message includes the reporting time of the corresponding image, an image description, a description of the behavioral event, and a description of the person. Each image is obtained by a smart entry device capturing images of behavioral events at the door. Specifically, in this application, the smart entry device captures images of behavioral events at the door, generates corresponding messages for the obtained images, and stores them in the message database. Upon obtaining a target user question related to entry, filtering conditions are determined. These filtering conditions include a time range based on the intent represented by the target user question. Candidate messages that meet the filtering conditions are filtered from the message database, achieving automatic filtering of behavioral events based on the target user question. Furthermore, this application can rewrite each candidate message obtained from the filtering to obtain each target message. Using a predetermined large model, feedback results matching the target user question are generated to achieve automatic analysis of the filtered behavioral events. As can be seen, this application, targeting user issues related to home entry (issues related to behavioral events at the door), can select candidate messages from a message database based on the given time range of the user issue, thus filtering out behavioral events matching the time range. Furthermore, it obtains the identities of the individuals involved in the behavioral events through message rewriting, and uses a large model based on the rewritten target message to generate feedback results for the user issue. Therefore, this application can automatically filter and analyze behavioral events, improving the intelligence level of behavioral event filtering and analysis.

[0022] Furthermore, message rewriting involves merging the identities of the same individuals indicated in the descriptions of various candidate messages and assigning a unique identity to each indicated individual. This prevents the same individual from being assigned different identities, which could lead to inaccurate feedback. The prompts constructed in this application include at least the target user question, each target message, and explanations for each target message. This ensures that the pre-defined large model can accurately understand the target user question and each target message, thereby analyzing each target message and accurately generating feedback results that match the target user question.

[0023] Of course, implementing any product or method of this application does not necessarily require achieving all of the advantages described above at the same time. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other embodiments can be obtained based on these drawings.

[0025] Figure 1 A flowchart illustrating an information analysis method provided in an embodiment of this application;

[0026] Figure 2 A schematic diagram illustrating the process of image acquisition for behavioral events at the door provided in this embodiment of the application;

[0027] Figure 3 A schematic diagram illustrating the process of generating a message corresponding to any image, as provided in an embodiment of this application;

[0028] Figure 4 A flowchart illustrating the process of assigning a person identity to each merged person in an embodiment of this application;

[0029] Figure 5 This is a schematic diagram of the structure of an information analysis device provided in an embodiment of this application;

[0030] Figure 6 This is a block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0031] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art based on this application are within the scope of protection of this application.

[0032] Existing smart entry systems often only record individual behavioral events when they detect them at the door, such as the time of the event and the triggering device. When users need answers to any entry-related questions, such as "Were there any strangers lingering at the door this afternoon and for how long?" or "How many times has the delivery person visited this week?", users often have to manually filter and analyze numerous behavioral events to find the answer.

[0033] Based on this, embodiments of this application provide an information analysis method, apparatus, electronic device, and storage medium to automatically filter and analyze behavioral events, thereby improving the intelligence level of filtering and analyzing behavioral events.

[0034] The following section first introduces an information analysis method provided in this application.

[0035] The information analysis method provided in this application can be applied to electronic devices, which can be smart entry-level devices or interactive devices capable of controlling smart entry-level devices. These interactive devices can be terminal devices such as mobile phones or computers. This application does not limit the specific form of the electronic device. The information analysis method provided in this application can be applied to scenarios where a user needs an answer to a user question related to entry into the home. This user question may include the type of behavioral event the user wants to know (entry into the home can be understood as knocking, loitering, opening the door, etc.) and the time range, etc. This application does not limit the user question related to entry into the home.

[0036] For example, the execution entity of the information analysis method provided in this application can be an information analysis terminal. The information analysis terminal can be installed on a smart entry-exit device, or an interactive device connected to the smart entry-exit device for controlling the smart entry-exit device. Specifically, the information analysis terminal can be a client with information filtering and analysis functions, capable of filtering and analyzing behavioral events at the door (which can be understood as information summarization), thereby obtaining answers to user questions related to entry. Alternatively, the information analysis terminal can also be understood as the smart entry-exit device itself, or the interactive device itself; this application does not limit the specific form of the information analysis terminal.

[0037] An information analysis method provided in this application embodiment may include the following steps:

[0038] In response to obtaining target user questions related to home visits, filtering criteria are determined; wherein, the filtering criteria include a time range obtained based on the intent represented by the target user questions;

[0039] From the message database, candidate messages that meet the filtering conditions are filtered; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door.

[0040] The selected candidate messages are rewritten to obtain target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person.

[0041] Construct prompt words that include at least the target user's question, each target message, and an explanation of each target message;

[0042] The prompt word is input into a predetermined large model, so that the large model responds to the prompt word and generates a feedback result that matches the target user's question.

[0043] This application provides an information analysis method that includes a message database. Each message in the message database corresponds to an image. Each message includes the reporting time of the corresponding image, an image description, a description of the behavioral event, and a description of the person. Each image is obtained by a smart entry device capturing images of behavioral events at the door. Specifically, in this application, the smart entry device captures images of behavioral events at the door, generates corresponding messages for the obtained images, and stores them in the message database. Upon obtaining a target user question related to entry, filtering conditions are determined. These filtering conditions include a time range based on the intent represented by the target user question. Candidate messages that meet the filtering conditions are filtered from the message database, achieving automatic filtering of behavioral events based on the target user question. Furthermore, this application can rewrite each candidate message obtained from the filtering to obtain each target message. Using a predetermined large model, feedback results matching the target user question are generated to achieve automatic analysis of the filtered behavioral events. As can be seen, this application, targeting user issues related to home entry (issues related to behavioral events at the door), can select candidate messages from a message database based on the given time range of the user issue, thus filtering out behavioral events matching the time range. Furthermore, it obtains the identities of the individuals involved in the behavioral events through message rewriting, and uses a large model based on the rewritten target message to generate feedback results for the user issue. Therefore, this application can automatically filter and analyze behavioral events, improving the intelligence level of behavioral event filtering and analysis.

[0044] Furthermore, message rewriting involves merging the identities of the same individuals indicated in the descriptions of various candidate messages and assigning a unique identity to each indicated individual. This prevents the same individual from being assigned different identities, which could lead to inaccurate feedback. The prompts constructed in this application include at least the target user question, each target message, and explanations for each target message. This ensures that the pre-defined large model can accurately understand the target user question and each target message, thereby analyzing each target message and accurately generating feedback results that match the target user question.

[0045] The following description, in conjunction with the accompanying drawings, provides an exemplary method for information analysis based on an embodiment of this application.

[0046] like Figure 1 As shown in the embodiments of this application, an information analysis method may include the following steps:

[0047] S101: In response to obtaining questions about target users related to home visits, determine the screening criteria;

[0048] The filtering criteria include a time range derived from the intent represented by the target user's question;

[0049] After obtaining the target user questions related to home visits, considering that the target user questions are usually related to behavioral events within a certain period of time, this application can determine the filtering conditions; wherein, the filtering conditions include a time range obtained based on the intent represented by the target user questions, for example, the filtering conditions can be filtering conditions with the time range as the condition content; and, the time range is obtained based on the intent represented by the target user questions.

[0050] In addition, target user questions are related to home entry. Home entry can be understood as behavioral events at the door (including: knocking, lingering, opening the door, prying open the door, etc.). Target user questions can also be understood as questions related to smart home entry devices, such as questions related to behavioral events recorded by smart home entry devices, etc.

[0051] For example, in one implementation, before determining the filtering conditions in response to obtaining a target user question related to home entry, the method further includes: receiving a user question from a user; invoking a predetermined large model to perform intent analysis on the received user question and obtain analysis results; wherein the intent analysis is used to analyze whether the received user question is related to home entry and to analyze the indicated time range; if the analysis results indicate that the received user question is related to home entry, it is determined that a target user question related to home entry has been obtained.

[0052] The step of determining the filtering criteria in response to obtaining target user questions related to home visits includes: determining the filtering criteria based on the time range in the analysis results in response to obtaining target user questions related to home visits.

[0053] Before determining the screening criteria, it can be determined whether the acquired user questions are target user questions related to home entry. If not, subsequent screening and analysis steps are unnecessary. This application can receive user questions and perform intent analysis on the user questions by calling a pre-defined large model to analyze whether the user questions are related to home entry and the indicated time range, obtaining analysis results. If the analysis results indicate that the user questions are related to home entry, it is determined that target user questions related to home entry have been acquired, and subsequent steps for screening and analysis can be performed. Specifically, if the analysis results contain intents such as opening the door, lingering, loitering, or knocking, it can be determined that the received user questions are related to home entry.

[0054] After obtaining the analysis results of the large model's intent analysis of user questions, if it is determined that a target user question related to home entry has been obtained, and the analysis results contain the time range indicated by the target user question after intent analysis, then the filtering conditions can be determined based on the time range in the analysis results. For example, the time range in the analysis results can be directly determined as the filtering conditions, or a portion of the time range in the analysis results can be selected to obtain the filtering conditions, etc.

[0055] For example, the target user question could be "How long did the stranger linger at my doorstep this afternoon?", and the defined filtering criteria could be: 12:00 noon to 18:00 today; or 12:00 noon to 14:00 today.

[0056] This application can use a pre-defined large model to perform intent analysis on received user questions, thereby accurately determining whether the user question is a target user question related to home visits based on the analysis results; and, if the user question is a target user question related to home visits, it can also quickly determine the screening criteria based on the time range in the analysis results.

[0057] It is important to emphasize that a large model refers to an artificial intelligence model trained on massive amounts of data and possessing an extremely large parameter scale (usually in the hundreds of millions or even trillions). Its core characteristic is that through powerful computing capabilities and large amounts of data, it learns the complex rules, knowledge, and patterns contained in the data, thereby possessing cross-domain general intelligence or powerful capabilities for specific scenarios. For example, a large model can be a chatbot, a text generation tool, a language translator, or a text analysis tool, etc. This application does not limit the specific form of the large model.

[0058] It should be emphasized that the determination of whether a target user question related to home visit has been obtained can also be implemented in other ways and is not limited to the above-mentioned determination methods. For example, after receiving a user question, it is possible to analyze whether the user question contains keywords related to home visit and keywords related to time range. If they are contained, it is determined that a target user question related to home visit has been obtained, and based on the time information represented by the time range-related keywords in the target user question, a time range that can be used to construct filtering conditions can be determined.

[0059] S102: Filter candidate messages that meet the filtering conditions from the message database;

[0060] Each message in the message database corresponds to an image. Each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person involved, obtained based on the analysis of the corresponding image. In other words, each message contains the reporting time, image description, behavioral event description, and person involved, and all the contents are related to the corresponding image. Each image is an image obtained by the smart home access device from image acquisition of behavioral events in front of the door.

[0061] In this application, for each image obtained by the smart entry device from the image acquisition of behavioral events in front of the door, a message is generated. The message contains a description of the image, a description of the behavioral events in the image, and a description of the people involved, that is, it contains information that provides a comprehensive description of the behavioral events in the image. Thus, candidate messages that meet the filtering criteria can be filtered from the message database for subsequent analysis.

[0062] Candidate messages can be understood as messages in the message database that match the time range included in the filtering conditions. That is, the reporting time of the corresponding image in each candidate message belongs to the time range included in the filtering conditions. In other words, each candidate message is a message within the time range that represents the intent of the target user's question asked by the user. During filtering, candidate messages whose reporting time meets the time range included in the filtering conditions can be selected from the message database.

[0063] For example, in one implementation, the smart entry device acquires images of behavioral events at the door by: waking up the camera in response to detecting a moving object; performing human detection on the image captured by the camera; and, if a human figure is detected, controlling the camera to generate a target video recording; wherein the target video recording is a recording of the time period during which a human figure is detected; extracting keyframes from the target video recording to obtain the extracted image; and, in response to triggering any of the specified entry events, controlling the camera to capture an image at the time of occurrence of that entry event.

[0064] In this application, each message in the message database corresponds to an image obtained by the smart entry device from event collection of behavior in front of the door. That is, each image is related to entry and each image is an image of the behavior event in front of the door. When the smart entry device is collecting images, if it detects a moving object (the sensors of the smart entry device can detect the presence of a moving object), it wakes up the camera and performs human detection on the image captured by the camera. If a human figure is detected, the camera is controlled to generate a target video recording. After obtaining the target video recording, in order to obtain images of the behavior event in front of the door, keyframes can be extracted from the target video recording to obtain the extracted images. Each extracted image can also be called a keyframe image.

[0065] The target video recording refers to the recording during the time period when a human figure is detected. Specifically, the smart entry-level device can record the video during the time period when a human figure is detected, including the start and end times of the human figure's presence, thus generating the target video recording. The specific generation method is similar to that of existing cameras or recording devices, and will not be elaborated upon here. Furthermore, for example, when extracting keyframes from the target video recording, keyframes can be extracted at equal intervals, such as every 0.1 seconds, or every two frames, etc.; or, keyframes can be extracted randomly. This application does not limit the keyframe extraction method. The number of keyframes extracted for any target video recording can be a target number, such as 5, 4, or 3, etc. If a target video recording is short, such as 2 seconds or 1 second, the number of keyframes extracted for that target video recording may be less than the target number. This application does not limit the number of keyframes extracted.

[0066] In addition, the smart entry device can detect whether any of the specified entry events has been triggered. If any entry event is triggered, the device controls the camera to capture an image at the moment the event occurred; this image can also be called a captured image. The specified entry events can include knocking, opening the door, or other events. The smart entry device can detect whether a specific entry event has been triggered using its sensors.

[0067] When a smart entry device captures images of behavioral events at the door, it can record the image capture time. Furthermore, for keyframe images extracted from the same target video recording, the behavioral event description of the image can be marked as a loitering event or a wandering event. For any of the specified entry events triggered, the behavioral event description of the image can be marked as the triggered entry event. For example, if the smart entry device detects a door opening event, it controls the camera to capture an image at the moment the door opening event occurs, and marks the behavioral event description of that image as a door opening event. The smart entry device can also report the capture time and behavioral event description of each captured image for subsequent image message generation. The images obtained by extracting keyframes from the target video recording may include images at the moment a specific entry event occurs; that is, the keyframe images extracted from a target video recording may include captured images at the moment a specific entry event occurs. For example, if a human figure appears during the target video recording period and a door opening event occurs, this application does not limit this.

[0068] It should be noted that, for example, the sensors on the smart entry device can detect whether any of the specified entry events has been triggered, such as opening the door, ringing the doorbell, and picking the lock. The smart entry device can detect the event through the corresponding sensor, control the camera to capture the image, and label the behavioral event description of the captured image according to the event type detected by the sensor. The smart entry device may include one or more sensors, and can detect multiple types of entry events through the same sensor, or multiple sensors may detect multiple types of entry events separately (e.g., each sensor detects one type of entry event). This application does not limit this.

[0069] As can be seen, in this application, the intelligent entry device can generate target video recordings for the time period when a human figure appears, and perform keyframe abstraction to obtain keyframe images, that is, to capture images of loitering events (or wandering events) in front of the door (which may also include captured images at the time of occurrence of specified entry events); and, when any of the specified entry events is triggered, the device can capture images at the time of occurrence of the entry event through the camera, that is, to capture images of knocking, opening, and knocking events in front of the door, excluding loitering events, thereby realizing image capture of various behavioral events, so as to generate corresponding messages for various behavioral events, and facilitate subsequent automatic filtering and analysis of messages through a message database to obtain feedback results on the target user's questions.

[0070] For example, in one implementation, the method for generating a message corresponding to any image includes:

[0071] The predefined large model is invoked to perform an image description on the image, resulting in an image description of the image; human figure matting is performed on the image to obtain at least one human figure image of the image; human figure feature extraction and attribute feature extraction are performed on each human figure image of the obtained image to obtain a person description of the image; wherein, the person description of each human figure image includes the human figure feature vector and person attributes of the human figure image; based on the reporting time of the image, the image description, the behavioral event description, and the person description of each human figure image of the image, a message is constructed to obtain the message corresponding to the image; wherein, the reporting time and behavioral event description of each image are reported by the smart home access device.

[0072] Each message in this application includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person description obtained based on the analysis of the corresponding image. The smart home access device can report the reporting time (i.e., the acquisition time) and the description of the behavioral event for each image. The information analysis end can call a predetermined large model to perform image description on the image, and perform human figure matting on the image to obtain at least one human figure image of the image; and perform human figure feature extraction and attribute feature extraction on each human figure image of the image to obtain the person description of each human figure image in the image; then, based on the image description obtained by calling the large model, the reporting time and description of the behavioral event of the image reported by the smart home access device, and the obtained person description, a message is constructed to obtain the message corresponding to the image.

[0073] Image description refers to the comprehensive, accurate, and organized interpretation and presentation of the visual information contained in an image through language (such as text); that is, converting the visually visible elements (such as objects, people, scenes, colors, composition, etc.) and potential meanings (such as emotions, atmosphere, logical relationships, etc.) in the image into understandable language. Each behavioral event can have a flag bit. For example, for a captured image that was captured when a door opening event was triggered, the event type in the behavioral event description of the captured image can be a door opening event, and the flag bit is 1 (1 indicates that the event was triggered, 0 indicates that the event was not triggered). The description of each human figure image includes a human feature vector and human attributes. For example, the human figure image can be input into a human feature extraction model to obtain its human feature vector; and it can be input into a human attribute model (a model for extracting human attributes) to obtain its human attributes. For example, human attributes may include: age (child, youth, adult, middle-aged, or elderly), gender (male or female), orientation (front, side, or back), special identity (security guard, delivery person, food delivery worker, etc.), and sharpness, human figure completeness, and posture (e.g., standing or squatting posture), etc. Sharpness and human figure completeness can be represented as percentages. The human feature extraction model and human attribute model involved in this application are similar to existing models for extracting human features and human attributes, and will not be elaborated upon here.

[0074] As can be seen, this application can obtain the image description of the image through a predetermined large model, and can obtain the human feature and attribute feature extraction of at least one human image obtained by human figure matting of the image, thereby obtaining the human figure description of each human image in the image. Combined with the reporting time of the image reported by the smart home device (reported immediately after collection, and the reporting time can also be understood as the image collection time) and the behavioral event description, it is possible to construct information that provides a comprehensive description of the behavioral events of the image.

[0075] Optionally, each message may also include: a description of the specified item;

[0076] For any given image, the methods for generating the message corresponding to that image also include:

[0077] Perform detection on the image for the specified item, and obtain the detection results;

[0078] The message corresponding to the image is constructed based on the image's reporting time, image description, behavioral event description, and the character description of each human figure in the image, including:

[0079] Based on the image's reporting time, image description, behavioral event description, detection results, and the character description of each human figure in the image, a message is constructed to obtain the message corresponding to the image.

[0080] Each message can also include a description of a specified item. Before generating a message for each image, the image can be used to detect the specified item (specifically, object detection algorithms can be used, which will not be elaborated here). The detection results are then combined to construct the message corresponding to the image. The specified item can be a package, takeout food, or hazardous materials, etc., to enable reminders or alarms for the specified item. Subsequent feedback results can include reminders or alarms for the specified item.

[0081] It should be noted that the above-described method of generating corresponding messages for each image collected by the smart entry device in response to behavioral events at the door is merely an example. The messages may also contain other descriptions of the images or the content in the images, and should not constitute a limitation on this application.

[0082] For example, the format of the message obtained in this application may be:

[0083]

[0084] Wherein, `report_time` is the image reporting time, and `YYYY-MM-DD HH:MM:SS` represents the specific reporting time (year-month-day hour:minute:second). `description` is the image description, where "xxxxx" contains the specific content of the image description. `message_event` is the description of the image's behavioral event, where "open_door" indicates the type of behavioral event, i.e., the current behavioral event is an "open door" event, which can correspond to a flag bit 0 or 1. "person" refers to the description of the people contained in the image, including the descriptions of people p1 and p2 (at this point, their identities are not explicitly stated, they are simply labeled in order). In each person's description, `box:{x,y,w,h}` is the human shape feature vector (containing the position (x,y) of the human box, as well as its width `w` and height `h`); `user_define_identity` represents the user-defined identity, i.e., the person's ID (the ID indicating the order in which they appear in the image); `reid_vector` is the person re-identification feature vector; the "xxx" after `reid_vector` indicates the specific content of the person re-identification feature vector; `human_attr` represents the person's attributes, including: age:xxx, gender:xxx, orientation:xxx, and the feature `identification:xxx` representing the person's special identity (such as security guard, deliveryman, courier, etc.) (this feature can be 0 if the person's identity is not explicitly stated). `fire_flag`: 1 indicates fire information, 1 is the flag corresponding to the fire, indicating that fire exists in the current image (0 indicates it does not exist). It should be emphasized that the parameter content "xxx" or the flag bit corresponding to each parameter can be blank, and this application does not limit this.

[0085] S103: Rewrite the candidate messages obtained from the screening to obtain the target messages;

[0086] The message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person.

[0087] After filtering, in order to facilitate subsequent analysis of the messages, this application can rewrite the messages of each candidate message obtained by filtering. That is, for the same person indicated by the person description in each candidate message, the identities are merged and the indicated person is assigned a person identity to obtain each target message.

[0088] It should be noted that the acquisition time of the image corresponding to each candidate message is within the time range represented by the target user's question. That is, message rewriting is: merging the identities of the same person in each candidate message within the time range represented by the target user's question, and assigning the person an identity to obtain each target message.

[0089] For example, in one implementation, the step of rewriting the selected candidate messages to obtain target messages includes: based on the person descriptions in the selected candidate messages, merging the identities of the same persons indicated by the person descriptions in the candidate messages according to a predetermined merging strategy; assigning a person identity to each merged person; and rewriting the assigned person identity into the corresponding candidate message in each candidate message to obtain target messages.

[0090] It should be noted that after the smart home access device collects images of behavioral events at the door, it does not identify the person's identity. It can be understood that it only records the ID (identifier) ​​of the order in which the person appears in the image without comparison. In addition, the identity of the person contained in each message in the message database is not clearly defined.

[0091] When rewriting messages, this application can merge the identities of identical individuals indicated by the character descriptions in each candidate message according to a predetermined merging strategy. Each merged individual is then assigned a character identity and rewritten into the corresponding candidate message to obtain target messages. The characters contained in the target messages are those whose identities have been merged and clearly defined. Different target messages may contain individuals with the same character identity.

[0092] This application rewrites each candidate message, enabling the identification of the same person and explicitly assigning that person's identity. This avoids the situation where the same person is identified as multiple different people, and improves the accuracy of the feedback results obtained from subsequent analysis of each target message.

[0093] Optionally, the person description in each message includes: a human feature vector and person attributes; the predetermined merging strategy includes at least one of the following strategies: if the similarity of the human feature vectors in different person descriptions reaches a predetermined threshold, the identities indicated by the different person descriptions are merged; if the gender attributes contained in the person attributes in different person descriptions are different, or the age attributes do not match, the identities indicated by the different person descriptions are not merged.

[0094] Each message contains a description of a person, including a human feature vector and the person's attributes. When merging identities, if the similarity of the human feature vectors in different person descriptions reaches a predetermined threshold (such as 80%, 85%, or 90%), the identities indicated by the different person descriptions are merged. For example, if the similarity of the human feature vectors of Stranger_001 in one candidate message and Stranger_002 in another candidate message reaches the predetermined threshold, then Stranger_001 in one candidate message and Stranger_002 in another candidate message are merged, and the merged identity can be Stranger A.

[0095] If the gender attribute in the character descriptions of different characters differs, or if the age attribute does not match, the identities indicated by the different character descriptions will not be merged to avoid merging different characters into the same character. The age attribute can include childhood, youth, adulthood, middle age, or old age (attribute ranges in order of age). An age attribute mismatch can be understood as: the age attribute ranges contained in the character attributes of different characters are not adjacent. For example, if one character description contains the age attribute of youth, and another character description contains the age attribute of adulthood, then their age attributes match (e.g., 19 years old may be judged as youth or adulthood); if one character description contains the age attribute of childhood, and another character description contains the age attribute of adulthood, then their age attributes do not match.

[0096] Furthermore, if the orientation attribute of a message's person attribute is "backwards," then the person indicated by the person description in that message will not be used as the basis for identity merging (i.e., other people cannot be treated as the same person, to avoid inaccurate person descriptions due to the "backwards" orientation attribute, leading to incorrect identity merging). For messages with the orientation attribute being "front" or "sideways," the person indicated by the person description in that message will be used as the basis for identity merging. For example, if the similarity between the human feature vector in the person description of another message and the human feature vector in the person description of this message reaches a predetermined threshold, then the person indicated by the person description in the other message will be treated as the same person indicated by the person description of this message.

[0097] It should be noted that for two different character descriptions, if the gender attribute or age attribute contained in the humanoid attributes of the two character descriptions is different, even if the similarity of the humanoid feature vectors in the two character descriptions reaches a predetermined threshold, the identities indicated by the two different character descriptions cannot be merged. That is, the priority of the merging strategy for different gender attributes or mismatched age attributes is higher than the priority of the merging strategy for humanoid feature vector similarity reaching a predetermined threshold. In this way, it is possible to avoid erroneously merging people of different genders or different ages into the same person.

[0098] This application, through a predetermined merging strategy, can merge the identities indicated by different character descriptions and avoid erroneous identity merging, thereby improving the accuracy of identity merging.

[0099] It should be noted that the specific process of identity merging and assigning identities to individuals will be described in detail in subsequent embodiments, and will not be repeated here.

[0100] For example, a target message after message rewriting in this application can be:

[0101]

[0102] In the rewritten target message, the human-shaped feature vector box:{x,y,w,h} is omitted, and the meanings of the other parameters are similar to those in the message of the image obtained above, and will not be elaborated here.

[0103] S104: Construct prompt words that include at least the target user's question, each target message, and an explanation of each target message;

[0104] In order to fully understand each target message and analyze each target message in conjunction with the target user's question, this application constructs a prompt word that includes at least the target user's question, each target message, and an explanation of each target message, so that a predetermined large model can be invoked for analysis based on the prompt word.

[0105] For example, in one implementation, the prompt words constructed by this application can be:

[0106]

[0107]

[0108] Alternatively, the prompts may contain only the target user's question and an explanation of each target message, or only the explanation of each target message. When inputting into a predefined large model, the prompts and each target message may be input simultaneously, or the prompts, each target message, and the target user's question may be input simultaneously. This application does not limit this.

[0109] S105: Input the prompt word into a predetermined large model so that the large model responds to the prompt word and generates a feedback result that matches the target user's question.

[0110] After the prompt words are generated, this application can input the prompt words into a predefined large model. The large model analyzes each target message based on the prompt words and generates feedback results that match the target user's question.

[0111] For example, based on the above prompt words and three target messages, if the target user's question is "How long did a stranger linger at my door this afternoon?", the large model will determine the answer based on the prompt word "f", where 'user_define_identity' starting with "stranger" indicates a stranger, and the rest are acquaintances. In target message 3, "user_define_identity" is "father," indicating an acquaintance, while in target messages 1 and 2, it's all stranger A. The target user's question is about how long the lingering was, and the large model will calculate based on the time difference in "report_time" to arrive at the answer: A stranger A lingered for two minutes this afternoon. The following is the corresponding video recording: xxx (i.e., the target video recordings to which the images in target messages 1 and 2 belong; of course, the corresponding images can also be displayed directly, etc.).

[0112] In this application, after inputting prompt words, the large model automatically provides answers based on each target message, the target user's question, and the prompt words. Since the prompt words contain explanations of each target message (i.e., structural annotations of the message), the large model can accurately answer the target user's question and obtain feedback results.

[0113] In the technical solution of this application, all operations such as acquiring, storing, using, processing, transmitting, providing, and disclosing the target user problem, each message in the message database, an image corresponding to each message in the message database, person identity, image description, facial feature vector, human shape feature vector, and person attributes are carried out with the user's authorization.

[0114] It should be noted that the facial feature vector, human shape feature vector, and person attributes in this embodiment are not specific to any particular user and do not reflect the personal information of any particular user.

[0115] This application provides an information analysis method that includes a message database. Each message in the message database corresponds to an image. Each message includes the reporting time of the corresponding image, an image description, a description of the behavioral event, and a description of the person. Each image is obtained by a smart entry device capturing images of behavioral events at the door. Specifically, in this application, the smart entry device captures images of behavioral events at the door, generates corresponding messages for the obtained images, and stores them in the message database. Upon obtaining a target user question related to entry, filtering conditions are determined. These filtering conditions include a time range based on the intent represented by the target user question. Candidate messages that meet the filtering conditions are filtered from the message database, achieving automatic filtering of behavioral events based on the target user question. Furthermore, this application can rewrite each candidate message obtained from the filtering to obtain each target message. Using a predetermined large model, feedback results matching the target user question are generated to achieve automatic analysis of the filtered behavioral events. As can be seen, this application, targeting user issues related to home entry (issues related to behavioral events at the door), can select candidate messages from a message database based on the given time range of the user issue, thus filtering out behavioral events matching the time range. Furthermore, it obtains the identities of the individuals involved in the behavioral events through message rewriting, and uses a large model based on the rewritten target message to generate feedback results for the user issue. Therefore, this application can automatically filter and analyze behavioral events, improving the intelligence level of behavioral event filtering and analysis.

[0116] Furthermore, message rewriting involves merging the identities of the same individuals indicated in the descriptions of various candidate messages and assigning a unique identity to each indicated individual. This prevents the same individual from being assigned different identities, which could lead to inaccurate feedback. The prompts constructed in this application include at least the target user question, each target message, and explanations for each target message. This ensures that the pre-defined large model can accurately understand the target user question and each target message, thereby analyzing each target message and accurately generating feedback results that match the target user question.

[0117] Optionally, in another embodiment of this application, the identity of the person includes both acquaintances and strangers;

[0118] Assigning a character identity to each character after the merger includes:

[0119] For each merged person, the similarity between the person's human feature vector and the human feature vectors of each acquaintance in the first feature template library is calculated to obtain the calculation result.

[0120] If the calculation result indicates that the human feature vector of the person matches the human feature vector of an acquaintance, then the person's identity is assigned to the acquaintance identity of the matching acquaintance.

[0121] If the calculation result indicates that the human feature vector of this person does not match the human feature vectors of any acquaintances in the first feature template library, analyze whether there is a vector in the second feature template library that matches the human feature vector of this person.

[0122] If the character does not exist, assign the character a stranger identity, and if the character's attributes meet predetermined conditions, add the character's humanoid feature vector to the second feature module library; wherein, the predetermined conditions are conditions concerning orientation, clarity, and / or humanoid completeness.

[0123] If it exists, assign the character's identity to the character indicated by the matching vector;

[0124] The second feature template library is used to store: human feature vectors in each candidate message that do not match the human feature vectors of each acquaintance in the first feature template library and meet predetermined conditions, and the human attribute represented by the human feature vectors stored in the second feature template library meets predetermined conditions; after receiving the feedback result of the target user's question, the human feature vectors included in the second feature template library are cleared.

[0125] In this application, the identities of the people involved include acquaintances and strangers. Acquaintances can include: father, mother, grandfather, grandmother, etc., while strangers can be assigned identities in sequence, such as: stranger A, stranger B, stranger C, stranger D, etc. Furthermore, users can pre-define the identities of each acquaintance (for example, users can input images of each acquaintance through the user interface and label the identity of each acquaintance corresponding to the image). This application can extract features from the human-shaped images of each acquaintance and store the resulting human-shaped feature vectors in a first feature template library. The first feature template library can be understood as a template library used to store the human-shaped feature vectors of each acquaintance; and the first feature template library can contain multiple features with the same acquaintance identity, for example, containing multiple human-shaped feature vectors related to the father. This application does not limit this. The second feature template library is used to store human feature vectors in each candidate message that do not match the human feature vectors of each acquaintance in the first feature template library, and the human feature vectors stored in the second feature template library represent human attributes that meet predetermined conditions. After receiving feedback from the target user's question, the human feature vectors included in the second feature template library are cleared. That is, the second feature template library is used to temporarily store human feature vectors of strangers whose human attributes meet predetermined conditions, so that the human feature vectors of each person in the subsequent merging that match the human feature vector can be identified as the same stranger and assigned the same stranger identity, avoiding the incorrect identity assignment caused by the same stranger being assigned multiple different stranger identities in different messages.

[0126] The predetermined conditions are those concerning orientation, clarity, and / or human figure completeness. For example, the predetermined conditions are that the orientation is frontal or side, the clarity requirement is 80% (i.e., the human figure must be clear), and the human figure completeness requirement is 95% (i.e., the human figure must be complete). In order to use the human figure feature vector of the character that meets the predetermined conditions as the basis for identity assignment, the accuracy of identity assignment will be improved.

[0127] This application can calculate the similarity between the human feature vector of each merged person (if there are multiple human feature vectors for the person, one human feature vector can be selected for calculation) and the human feature vectors of each acquaintance in the first feature template library, and obtain the calculation result; if the human feature vector of the person matches the human feature vector of an acquaintance, then the person's identity can be assigned as the acquaintance identity of the matching acquaintance.

[0128] If a person's human feature vector does not match any of the human feature vectors of acquaintances in the first feature template library (meaning the person is not an acquaintance), then the system analyzes whether a vector (i.e., a stranger's human feature vector) matches the person's human feature vector in the second feature template library. If no such vector exists, the person is assigned a stranger identity, and if the person's attributes meet predetermined conditions, the person's human feature vector is added to the second feature template library (i.e., if a subsequent person's human feature vector matches this vector, then that subsequent person is assigned the same stranger identity as this human feature vector). If a match exists, meaning the person's human feature vector matches a stranger's human feature vector in the second feature template library, then the person's identity is assigned to the stranger identity indicated by the matching human feature vector.

[0129] In addition, the first feature template library and the second feature template library can also be integrated into the same database, and the first feature template library and the second feature template library belong to different storage areas of the same database, which is not limited in this application.

[0130] This application can use a first feature template library that stores the human feature vectors of acquaintances and a second feature template library that temporarily stores the human feature vectors of strangers to calculate the similarity of the human feature vectors of each merged person, thereby accurately assigning a person identity to each merged person. Furthermore, the second feature template library can also avoid the situation where the same stranger is assigned multiple different stranger identities in different messages, which would lead to incorrect identity assignment.

[0131] Optionally, the description of the person in each message may also include: a facial feature vector;

[0132] Assigning a character identity to each of the merged characters also includes:

[0133] Before analyzing whether there is a vector in the second feature template library that matches the human feature vector of the person, identify whether the message to which the person's description belongs corresponds to a captured image;

[0134] If so, the facial feature vector of the person is compared with each specified facial feature vector to calculate the similarity. If the calculated result indicates that the facial feature vector of the person matches a specified facial feature vector, then the person's identity is assigned to the acquaintance identity represented by the specified facial feature vector. Otherwise, the step of analyzing whether there is a vector in the second feature template library that matches the person's human-shaped feature vector is triggered. Here, each specified facial vector is the facial feature vector of each acquaintance recorded in the smart home access device.

[0135] If not, trigger the step of analyzing whether there is a vector in the second feature template library that matches the humanoid feature vector of the person.

[0136] In captured images, for example, images capturing a door opening event, the human figure may not be complete. This could lead to a situation where, even if the person is an acquaintance, they might be identified as a stranger based on the human figure's feature vector. In this application, the person description in each message may also include a facial feature vector. When assigning a person identity to each merged person, the facial feature vector can be used to correct the stranger identity assigned to the human figure in the captured image.

[0137] For any merged person, before analyzing whether there is a matching feature vector in the second feature template library, it is determined whether the message to which the person's description belongs corresponds to a captured image. If not, the person's feature vector is considered to be obtained from a complete human figure, triggering the step of analyzing whether there is a matching vector in the second feature template library. If so, the similarity between the person's facial feature vector and the specified facial feature vectors of each acquaintance is calculated. If the person's facial feature vector matches a specified facial feature vector, the person's identity is assigned to the acquaintance identity represented by the specified facial feature vector. That is, although the human feature vector does not match the human feature vectors of each acquaintance in the first feature template library, the facial feature matches a specified facial feature of an acquaintance, so the merged person is assigned the acquaintance identity represented by the specified facial feature. Otherwise, even if the person is identified as a stranger through facial features, the step of analyzing whether there is a matching vector in the second feature template library is triggered, and the merged person is assigned the corresponding stranger identity.

[0138] It should be noted that the specified facial vector refers to the facial feature vectors of various acquaintances recorded in the smart home access device. For example, the smart home access device can pre-record the facial feature vectors of multiple acquaintances: Acquaintance 1, Acquaintance 2, Acquaintance 3, and Acquaintance 4. The acquaintances recorded in the smart home access device are merely their IDs and do not specify their specific identities. Based on the correspondence between the facial feature vectors of each acquaintance recorded in the smart home access device and their corresponding identities, the identity of the person can be assigned to the acquaintance identity represented by the specified facial feature vector. The correspondence between the facial feature vectors of each acquaintance and their corresponding identities can be pre-defined by the user. For example, Acquaintance 1 corresponds to the identity of the father, Acquaintance 2 corresponds to the identity of the mother, Acquaintance 3 corresponds to the identity of the grandfather, and Acquaintance 4 corresponds to the identity of the grandmother. This application does not limit this.

[0139] As can be seen, when assigning a person identity to each merged person, this application can also correct the identity of people identified as strangers by facial feature vectors, based on human shape feature vectors, to avoid misidentifying familiar people as strangers in scenarios such as captured images by relying solely on human shape feature vectors.

[0140] The following describes an information analysis method provided in this application, with reference to another embodiment.

[0141] like Figure 2 As shown, the process of intelligent entry devices capturing images of behavioral events at the door can include the following steps:

[0142] S201: Human detection; that is, when the sensors (such as infrared sensors) of the smart home device detect the presence of a moving object, the camera is activated, the data stream of the image captured by the camera is obtained, and human detection is performed on the image captured by the camera.

[0143] S202: Determine if there is a person; that is, determine if a human figure is detected. If yes, proceed to step S203; otherwise, return to step S201 to continue human figure detection.

[0144] S203: Generate target video; that is, if a human figure is detected, a target video will be generated during the time period in which the human figure appears.

[0145] S204: Keyframe extraction; After obtaining the target video, abstract keyframes at equal intervals according to the time sequence. For example, the five keyframe images extracted are p1-p5.

[0146] S205: Home entry event detection; that is, after obtaining the data stream, detect whether any of the specified home entry events has occurred.

[0147] S206: Determine if there is a person; that is, after detecting any entry event, determine if there is a human figure. If yes, proceed to step S207; otherwise, return to step S205 to continue detecting entry events.

[0148] S207: Generate a captured image; that is, if an entry event occurs and a human figure is detected, the entry event flag is returned. For example, if a door opening event occurs, the flag is returned as 1, and an image pe is captured at the moment the entry event occurs.

[0149] It should be noted that the above steps S201-S207 are performed by the smart entry device. At this time, the smart entry device can act as an event monitoring module to monitor and capture images of behavioral events in front of the door. Furthermore, when the smart entry device is started, it can continuously perform the above steps to capture images.

[0150] like Figure 3As shown, the process of generating a message corresponding to any image may include the following steps:

[0151] S301: Call the large model to perform image description; that is, for any image among the keyframe images p1-p5 and the captured image pe, call the large model to perform image content description on the image to obtain the image description of the image.

[0152] S302: Human figure matting; that is, performing human figure matting on any image to obtain at least one human figure image of the image; and performing human figure feature extraction and attribute feature extraction on each human figure image of the obtained image, for example: returning human figure feature vector through human figure feature extraction model; and obtaining human attributes through human figure attribute model, including age (child, youth, adult, middle-aged or elderly), gender (male or female), orientation (front, side or back), special identity (security guard, courier, food delivery), etc., thereby obtaining the human figure feature vector and human attributes of the image.

[0153] S303: Construct a message; that is, combine the image description obtained through the large model, the human feature vector and human attributes of the image obtained by human figure matting, and the behavior event description and the image reporting time (the behavior event description and the image reporting time can be obtained by the smart home device) to construct a message for the image. The message format of each image is the same, see the example in step S102 above.

[0154] Additionally, each image can be used to detect specified items, such as takeout food, express delivery, and fireworks, and the detection results can be added to the message. It should be noted that steps S301-S303 above can be understood as being executed by the message generation module. This module can be located in the smart home device or be a functional module of the information analysis end. Whenever the message generation module acquires an image, it generates the corresponding message and stores it in the message database.

[0155] Before rewriting the message, we first collect the user question q1 (question text) and perform intent analysis through a large model. If we determine that the user question is a target user question related to home visits, we determine the filtering conditions. For example, if the target user question is "What time did Dad come home last night?", then the filtering condition r1 is {question_flg:1,start_time:%Y-%m-%d%H:%M:%S,end_time:%Y-%m-%d%H:%M:%S}.

[0156] Based on the filtering criteria, n candidate messages (x1-xn) are obtained from the message database. The n messages are then rewritten: First, the identities of the same person indicated by the person description in the n candidate messages are merged according to the predetermined merging strategy. If the orientation attribute in the person description is "back", then the person's human feature vector is not merged (or is not used as the merging basis). After merging, each person is assigned a person identity.

[0157] like Figure 4 As shown, the process of assigning a character identity to each merged character may include the following steps:

[0158] S401: Similarity calculation; For a human feature vector after identity merging, calculate the similarity with the human feature vectors of each acquaintance in the first feature template library, and obtain the calculation result.

[0159] S402: Does the similarity match an acquaintance? Based on the similarity calculation result, determine whether the similarity matches an acquaintance in the first feature template library. If yes, proceed to step S404; otherwise, proceed to step S403.

[0160] S403: Does the second feature template library contain a matching vector? That is, the human-shaped feature vector after identity merging does not match any of the acquaintances in the first feature template library. Continue to determine whether the second feature template library contains a matching vector. If yes, proceed to step S405; otherwise, proceed to step S406.

[0161] S404: Assign acquaintance identity; that is, if the similarity matches an acquaintance in the first feature template library, then the person is assigned the acquaintance identity of the matched acquaintance.

[0162] S405: Assign a stranger identity; that is, if a matching vector exists in the second feature template library, then the person is assigned the stranger identity indicated by the matching vector.

[0163] S406: Add to the second feature template library; that is, if there is no matching vector in the second feature template library, the person is assigned the identity of a stranger, and if the person's attributes meet the predetermined conditions, the person's humanoid feature vector is added to the second feature template library.

[0164] Furthermore, for individuals in the captured image, considering both facial features and human shape feature vectors (the face bounding box is located within the human shape bounding box), if the human shape feature vector identifies the individual as a stranger (not matching the human shape feature vectors of any acquaintances in the first feature template library), the individual is not initially assigned a stranger identity. Instead, the facial features are used to determine if the individual is an acquaintance. If the facial features identify the individual as an acquaintance, then the individual's identity is assigned as that of an acquaintance whose facial features match. Afterward, the assigned individual identity is rewritten into n candidate messages to obtain the target messages. The rewritten target messages are illustrated in the example in step S103 and will not be elaborated upon here.

[0165] To obtain feedback results tailored to the target user's question, it is necessary to construct prompt words. These prompt words should at least contain an explanation of the message (used to guide the large model in understanding the message format). They should be used in conjunction with the in-home scenario to help the large model understand the message format and obtain more accurate feedback results. Then, the rewritten n target messages, the current time t, the target user question q1, and the constructed prompt words are input into the large model. The large model understands the target messages and the target user question, and, combined with the current time t, outputs feedback results that match the target user question.

[0166] This application dynamically summarizes entry events based on user questions, rather than directly providing descriptions (event descriptions, image descriptions, and person descriptions, etc.) for users to analyze themselves. It automatically filters and analyzes behavioral events, improving the intelligence of this filtering and analysis process. Based on the entry scenario, this application designs differentiated prior information: entry events (opening the door, prying open the door, knocking on the door), and the identities of special strangers (delivery, takeout, etc.). This information is used for message generation and rewriting, resulting in accurate, filtered target messages. Furthermore, the prior information and message structure of this application are designed for message rewriting, producing comprehensive, filtered target messages that allow for full analysis by a large model, yielding comprehensive feedback results.

[0167] Based on the above method embodiments, this application also provides an information analysis device, such as... Figure 5 As shown, the device includes:

[0168] The determining module 510 is configured to determine filtering conditions in response to obtaining target user questions related to home visits; wherein the filtering conditions include a time range obtained based on the intent represented by the target user questions;

[0169] The filtering module 520 is used to filter candidate messages that meet the filtering conditions from the message database; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door.

[0170] The rewriting module 530 is used to rewrite the candidate messages obtained from the screening to obtain the target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in the candidate messages and assign a person identity to the indicated person.

[0171] The construction module 540 is used to construct prompt words that include at least the target user's question, the various target messages, and the explanation content for the various target messages;

[0172] Input module 550 is used to input the prompt word into a predetermined large model, so that the large model responds to the prompt word and generates a feedback result that matches the target user's question.

[0173] Optionally, the device further includes a judgment module, configured to: receive a user question from a user; invoke a predetermined large model to perform intent analysis on the received user question and obtain analysis results; wherein the intent analysis is used to analyze whether the received user question is related to home visit and to analyze the indicated time range; if the analysis results indicate that the received user question is related to home visit, determine that a target user question related to home visit has been obtained.

[0174] The determining module is specifically used to: in response to obtaining target user questions related to home visits, determine screening conditions based on the time range in the analysis results.

[0175] Optionally, the intelligent entry device acquires images of behavioral events in front of the door in the following ways: in response to detecting a moving object, the camera is woken up; human detection is performed on the image captured by the camera; if a human figure is detected, the camera is controlled to generate a target video recording; wherein, the target video recording is the recording during the time period in which the human figure is detected; keyframes are extracted from the target video recording to obtain the extracted image.

[0176] In addition, in response to triggering any of the specified entry events, the camera is controlled to capture an image at the moment the entry event occurs.

[0177] Optionally, for any given image, the method for generating a message corresponding to that image includes: calling the predetermined large model to perform an image description on the image, obtaining an image description of the image; performing human figure matting on the image to obtain at least one human figure image of the image; performing human figure feature extraction and attribute feature extraction on each of the obtained human figure images to obtain a person description of the image for each human figure image; wherein, the person description of each human figure image includes the human figure feature vector and person attributes of the human figure image; constructing a message based on the reporting time, image description, behavioral event description of the image, and the person description of each human figure image of the image to obtain a message corresponding to the image; wherein, the reporting time and behavioral event description of each image are reported by the smart home access device.

[0178] Optionally, each message may further include: a description of the specified item; and for any image, the method of generating the message corresponding to the image may further include: performing detection on the image regarding the specified item to obtain the detection result;

[0179] The message corresponding to the image is constructed based on the image's reporting time, image description, behavioral event description, and the character description of each human figure in the image. This includes: constructing a message based on the image's reporting time, image description, behavioral event description, detection result, and the character description of each human figure in the image to obtain the message corresponding to the image.

[0180] Optionally, the rewriting module includes:

[0181] The merging submodule is used to merge the identities of the same person indicated by the person descriptions in the various candidate messages obtained through filtering, according to a predetermined merging strategy.

[0182] Assign a submodule to assign a character identity to each character after merging;

[0183] The rewrite submodule is used to rewrite the assigned character identity into the corresponding candidate message in each candidate message to obtain each target message.

[0184] Optionally, the person description in each message includes: a human feature vector and person attributes; the predetermined merging strategy includes at least one of the following strategies: if the similarity of the human feature vectors in different person descriptions reaches a predetermined threshold, the identities indicated by the different person descriptions are merged; if the gender attributes contained in the person attributes in different person descriptions are different, or the age attributes do not match, the identities indicated by the different person descriptions are not merged.

[0185] Optionally, the identity of the person includes both acquaintances and strangers; the assigning submodule is used for:

[0186] For each merged person, the similarity between the person's human feature vector and the human feature vectors of each acquaintance in the first feature template library is calculated to obtain the calculation result.

[0187] If the calculation result indicates that the human feature vector of the person matches the human feature vector of an acquaintance, then the person's identity is assigned to the acquaintance identity of the matching acquaintance.

[0188] If the calculation result indicates that the human feature vector of this person does not match the human feature vectors of any acquaintances in the first feature template library, analyze whether there is a vector in the second feature template library that matches the human feature vector of this person.

[0189] If the character does not exist, assign the character a stranger identity, and if the character's attributes meet predetermined conditions, add the character's humanoid feature vector to the second feature module library; wherein, the predetermined conditions are conditions concerning orientation, clarity, and / or humanoid completeness.

[0190] If it exists, assign the character's identity to the character indicated by the matching vector;

[0191] The second feature template library is used to store: human feature vectors in each candidate message that do not match the human feature vectors of each acquaintance in the first feature template library, and the human attributes represented by the human feature vectors stored in the second feature template library meet predetermined conditions; after receiving the feedback result of the target user's question, the human feature vectors included in the second feature template library are cleared.

[0192] Optionally, the description of the person in each message may also include: a facial feature vector;

[0193] The assigned submodule is also used for:

[0194] Before analyzing whether there is a vector in the second feature template library that matches the human feature vector of the person, identify whether the message to which the person's description belongs corresponds to a captured image;

[0195] If so, the facial feature vector of the person is compared with each specified facial feature vector to calculate the similarity. If the calculated result indicates that the facial feature vector of the person matches a specified facial feature vector, then the person's identity is assigned to the acquaintance identity represented by the specified facial feature vector. Otherwise, the step of analyzing whether there is a vector in the second feature template library that matches the person's human-shaped feature vector is triggered. Here, each specified facial vector is the facial feature vector of each acquaintance recorded in the smart home access device.

[0196] If not, trigger the step of analyzing whether there is a vector in the second feature template library that matches the humanoid feature vector of the person.

[0197] This application also provides an electronic device, such as... Figure 6 As shown, it includes:

[0198] Memory 601 is used to store computer programs;

[0199] The processor 602, when executing the program stored in the memory 601, implements the steps of any of the described information analysis methods.

[0200] Furthermore, the aforementioned electronic device may also include a communication bus and / or a communication interface, with the processor 602, communication interface, and memory 601 communicating with each other via the communication bus.

[0201] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0202] The communication interface is used for communication between the aforementioned electronic devices and other devices.

[0203] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0204] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0205] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of any of the above-described information analysis methods.

[0206] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the information analysis methods described in the above embodiments.

[0207] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a solid-state drive (SSD), etc.

[0208] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0209] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the apparatus embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0210] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. An information analysis method, characterized in that, The method includes: In response to obtaining target user questions related to home visits, filtering criteria are determined; wherein, the filtering criteria include a time range obtained based on the intent represented by the target user questions; From the message database, candidate messages that meet the filtering conditions are filtered; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door. The selected candidate messages are rewritten to obtain target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person. Construct prompt words that include at least the target user's question, each target message, and an explanation of each target message; The prompt word is input into a predetermined large model, so that the large model responds to the prompt word and generates a feedback result that matches the target user's question.

2. The method according to claim 1, characterized in that, Before determining the screening criteria after obtaining the target user questions related to home visits, the method further includes: Receive user questions; A predefined large model is invoked to perform intent analysis on the received user questions and obtain analysis results; wherein, the intent analysis is used to analyze whether the received user questions are related to home visits and the time range indicated by the analysis. If the analysis results indicate that the received user questions are related to home visits, it is determined that target user questions related to home visits have been obtained. The response to obtaining target user questions related to home visits, and the determination of screening criteria, include: In response to obtaining target user questions related to home visits, filtering criteria are determined based on the time range in the analysis results.

3. The method according to claim 1, characterized in that, The intelligent entry device collects images of behavioral events in front of the door in the following ways: The camera is activated in response to the detection of a moving object; The camera captures footage and performs human detection. If a human figure is detected, the camera is controlled to generate a target video recording. The target video recording is the recording of the time period during which a human figure is detected. Keyframes are extracted from the target video recording to obtain the extracted image; as well as, In response to any of the specified entry events, the camera is controlled to capture an image at the moment the entry event occurs.

4. The method according to any one of claims 1-3, characterized in that, For any given image, the methods for generating the message corresponding to that image include: The predetermined large model is invoked to perform an image description on the image, thereby obtaining the image description of the image; Perform human figure cutout on the image to obtain at least one human figure image of the image; For each human figure image obtained from the image, human figure features and attribute features are extracted to obtain a description of the person in the image; wherein, the description of the person in each human figure image includes the human figure feature vector and the person's attributes. Based on the reporting time, image description, behavioral event description, and the character description of each human figure in the image, a message is constructed to obtain the message corresponding to the image; wherein, the reporting time and behavioral event description of each image are reported by the smart home access device.

5. The method according to claim 4, characterized in that, Each message also includes: a description of the specified item; For any given image, the methods for generating the message corresponding to that image also include: Perform detection on the image for the specified item, and obtain the detection results; The message corresponding to the image is constructed based on the image's reporting time, image description, behavioral event description, and the character description of each human figure in the image, including: Based on the image's reporting time, image description, behavioral event description, detection results, and the character description of each human figure in the image, a message is constructed to obtain the message corresponding to the image.

6. The method according to any one of claims 1-3, characterized in that, The process of rewriting each candidate message obtained from the screening to obtain each target message includes: Based on the character descriptions in each candidate message obtained through screening, the identities of the same characters indicated by the character descriptions in each candidate message are merged according to a predetermined merging strategy. Assign a character identity to each character after the merger; The assigned character identity is rewritten into the corresponding candidate message in each of the candidate messages to obtain each target message.

7. The method according to claim 6, characterized in that, The description of the person in each message includes: human-shaped feature vector and person attributes; The predetermined merging strategy includes at least one of the following strategies: If the similarity of the human feature vectors in different character descriptions reaches a predetermined threshold, the identities indicated by the different character descriptions will be merged. If the gender attribute in different character descriptions is different, or the age attribute is mismatched, then the identities indicated by the different character descriptions will not be merged.

8. The method according to claim 7, characterized in that, The identities of the people mentioned include both acquaintances and strangers; Assigning a character identity to each character after the merger includes: For each merged person, the similarity between the person's human feature vector and the human feature vectors of each acquaintance in the first feature template library is calculated to obtain the calculation result. If the calculation result indicates that the human feature vector of the person matches the human feature vector of an acquaintance, then the person's identity is assigned to the acquaintance identity of the matching acquaintance. If the calculation result indicates that the human feature vector of this person does not match the human feature vectors of any acquaintances in the first feature template library, analyze whether there is a vector in the second feature template library that matches the human feature vector of this person. If the character does not exist, assign the character a stranger identity, and if the character's attributes meet predetermined conditions, add the character's humanoid feature vector to the second feature module library; wherein, the predetermined conditions are conditions concerning orientation, clarity, and / or humanoid completeness. If it exists, assign the character's identity to the character indicated by the matching vector; The second feature template library is used to store: human feature vectors in each candidate message that do not match the human feature vectors of each acquaintance in the first feature template library, and the human attributes represented by the human feature vectors stored in the second feature template library meet predetermined conditions; after receiving the feedback result of the target user's question, the human feature vectors included in the second feature template library are cleared.

9. The method according to claim 8, characterized in that, The description of the person in each message also includes: facial feature vector; Assigning a character identity to each of the merged characters also includes: Before analyzing whether there is a vector in the second feature template library that matches the human feature vector of the person, identify whether the message to which the person's description belongs corresponds to a captured image; If so, the facial feature vector of the person is compared with each specified facial feature vector to calculate the similarity. If the calculated result indicates that the facial feature vector of the person matches a specified facial feature vector, then the person's identity is assigned to the acquaintance identity represented by the specified facial feature vector. Otherwise, the step of analyzing whether there is a vector in the second feature template library that matches the person's human-shaped feature vector is triggered. Here, each specified facial vector is the facial feature vector of each acquaintance recorded in the smart home access device. If not, trigger the step of analyzing whether there is a vector in the second feature template library that matches the humanoid feature vector of the person.

10. An information analysis device, characterized in that, The device includes: The determination module is used to determine filtering conditions in response to obtaining target user questions related to home visits; wherein the filtering conditions include a time range obtained based on the intent represented by the target user questions; The filtering module is used to filter candidate messages that meet the filtering conditions from the message database; wherein, each message in the message database corresponds to an image, and each message includes the reporting time and image description of the corresponding image, as well as a description of the behavioral event and the person based on the analysis of the corresponding image; each image is an image obtained by the smart home access device in response to the behavioral event at the door. The rewriting module is used to rewrite the candidate messages obtained from the screening to obtain the target messages; wherein, the message rewriting is used to merge the identities of the same person indicated by the person description in each candidate message, and to assign a person identity to the indicated person. A building module is used to build prompts that include at least the target user's question, the various target messages, and explanations for the various target messages; An input module is used to input the prompt words into a predetermined large model, so that the large model responds to the prompt words and generates a feedback result that matches the target user's question.

11. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method described in any one of claims 1-9.