Query result determination method and device, storage medium and electronic device

By separating and intelligently identifying the audio and video data collected by home smart cameras, generating feature text data, and adding priority tags to video clips, the problem of in-depth analysis of video content and accurate identification of feature tags in the existing technology is solved, and the satisfaction of user personalized information needs and optimization of viewing experience is achieved.

CN119917677APending Publication Date: 2025-05-02QINGDAO HAIER INTELLIGENT HOME APPLIANCE TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411843305.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art cannot deeply analyze video content and accurately identify feature tags, which cannot meet users' personalized information needs for monitoring objects.

Method used

By obtaining the original audio and video data collected by the target device, the separation process is performed to obtain multiple sub-audio and video data sets, they are inputted into the target model for identification, and text data for describing the object characteristics and changes are generated, and tags to be viewed for audio and video clips are added according to the monitoring requirements set by the management object.

Benefits of technology

It realizes in-depth analysis of video content and accurate identification of feature tags, meets users' personalized information needs for monitoring objects, and optimizes the viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119917677A_ABST
    Figure CN119917677A_ABST
Patent Text Reader

Abstract

The invention discloses a determination method and device of a query result, a storage medium and an electronic device, and relates to the technical field of smart home, the determination method of the query result comprises the following steps: obtaining original audio and video data collected by target equipment; separating the original audio and video data to obtain a plurality of sub-audio and video data sets, and inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, the target large model comprising a plurality of sub-models, and each sub-model corresponding to an object type; the sub-model is used for identifying object features in the sub-audio and video data set and generating text data for describing the object features and feature changes; and according to a monitoring demand set by the management object and the target identification result, adding a to-be-viewed tag to the audio and video segments in the original audio and video data, the to-be-viewed tag being used for managing the priority of viewing different audio and video segments in the original audio and video data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of smart home technology, and more specifically, to a method, device, storage medium and electronic device for determining a query result. Background Art

[0002] With the popularization of smart home cameras, different family members have more and more personalized demands for the content monitored by the camera, and such monitoring as parents' monitoring of the movements and what they are doing for their children and the elderly, and young people's monitoring of the movements of their pets will become the mainstream usage demand. At present, based on user settings, home cameras will report a large amount of captured video clips such as human shape, face recognition, screen changes, and pets. Based on the video content of human shape, face, and pet recognition, the video subject and its movements and what they are doing are characterized, screened, analyzed, and summarized based on the intelligent language model, and the associated data with clear feature labels are obtained, and targeted message notifications are made to users who have the need to view this part of personalized data. The data generated by existing video surveillance is a data display of a single subject, and there is no more detailed and concrete analysis and labeling of the subject appearing in the data.

[0003] Therefore, no effective solution has been proposed for the problem in related technologies that it is impossible to deeply analyze video content and accurately identify feature tags to meet users' personalized information needs for monitored objects. Summary of the invention

[0004] The embodiments of the present application provide a method, device, storage medium and electronic device for determining a query result, so as to at least solve the problem in the related art that it is impossible to deeply analyze video content and accurately identify feature tags to meet the user's personalized information needs for the monitored object.

[0005] According to an embodiment of the embodiments of the present application, a method for determining a query result is provided, comprising: obtaining original audio and video data collected by a target device; separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes; adding a to-be-viewed label to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition results, wherein the to-be-viewed label is used to manage the priority of viewing different audio and video segments in the original audio and video data.

[0006] In an exemplary embodiment, tags to be viewed are added to audio and video clips in the original audio and video data according to monitoring requirements set for the management object and target recognition results, including: parsing the monitoring requirements to obtain characteristics of the object to be monitored; wherein the object characteristics are used to indicate the type characteristics of the target object to be monitored and the specific behavior characteristics of the target object to be monitored; filtering the result content that meets the object characteristics from the target recognition results; filtering out multiple audio and video clips from the original audio and video data based on the correspondence between the result content and the original audio and video data, and adding viewing tags to the multiple audio and video clips respectively, wherein the viewing tags are used to identify the audio and video to be viewed that meet the monitoring requirements and appear in the original audio and video data.

[0007] In an exemplary embodiment, before adding viewing tags to multiple audio and video clips respectively, the above method also includes: determining the audio and video correlation between any two clips among the multiple audio and video clips; when the audio and video correlation is greater than or equal to a preset correlation, determining the associated audio and video clips as the same type of events to which the same viewing tag is to be added; when the audio and video correlation is less than the preset correlation, determining to add a different viewing tag for each audio and video clip.

[0008] In an exemplary embodiment, after adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition results, the above method also includes: determining the target attention features corresponding to the set monitoring requirements; labeling the target attention features to obtain multiple attention tags; matching the multiple attention tags with all the viewing tags that have been added; and determining a text message to be pushed to the management object of the target device based on the matching results, wherein the text message is used to carry a description text of the audio and video content that meets the set monitoring requirements.

[0009] In an exemplary embodiment, after adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements and target recognition results set by the management object, the above method also includes: when receiving an audio and video viewing request input by other objects, determining the target tag feature corresponding to the current audio and video viewing request; using the target tag feature to search in the target database, and synchronizing the search results to the target terminal corresponding to the other object for display, wherein the target database is used to store the original audio and video data to which the tags to be viewed have been added.

[0010] In an exemplary embodiment, before searching in the target database using the target tag feature, the method further includes: determining a first search permission for other objects corresponding to the audio and video viewing request; allowing other objects to perform real-time viewing of audio and video data when the first search permission is the same as a second search permission set for the management object; prohibiting other objects from performing real-time viewing of audio and video data when the first search permission is different from the second search permission set for the management object, and notifying the management object of the request data of other objects.

[0011] In an exemplary embodiment, before adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements and target recognition results set by the management object, the above method also includes: counting the changes in historical monitoring requirements; dynamically adjusting the generation rules of the tags to be viewed according to the changes in requirements, wherein the generation rules include at least: the type of tags to be viewed, the description of the tags to be viewed, and the identification criteria corresponding to the tags to be viewed; and updating the tag content of the tags to be viewed based on the generation rules.

[0012] According to another embodiment of the embodiment of the present application, a device for determining a query result is also provided, including: an acquisition module, used to acquire original audio and video data collected by a target device; a processing module, used to separate and process the original audio and video data to obtain multiple sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; an input module, used to input the multiple sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes multiple sub-models, each sub-model corresponds to an object type; the sub-model is used to identify object features in the sub-audio and video data set and generate text data for describing the object features and feature changes; an adding module, used to add a to-be-viewed label to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition results, wherein the to-be-viewed label is used to manage the priority of viewing different audio and video clips in the original audio and video data.

[0013] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned method for determining the query result when running.

[0014] According to another aspect of an embodiment of the present application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the method for determining the query result through the computer program.

[0015] According to another embodiment of the present application, a computer program product is provided, including a computer program, which implements the steps in any of the above query result determination method embodiments when the computer program is executed by a processor.

[0016] In an embodiment of the present application, the original audio and video data collected by the target device is obtained; the original audio and video data is separated and processed to obtain multiple sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; the multiple sub-audio and video data sets are input into the target large model to obtain the target recognition result, wherein the target large model includes multiple sub-models, each sub-model corresponds to an object type; the sub-model is used to identify the object features in the sub-audio and video data set and generate text data for describing the object features and feature changes; according to the monitoring requirements set by the management object and the target recognition result, the audio and video clips in the original audio and video data are added with a to-be-viewed tag, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data. The above technical solution solves the problem of not being able to deeply analyze the video content and accurately identify the feature tags to meet the user's personalized information needs for the monitored object. Furthermore, the audio and video data of multiple types of objects are analyzed by the intelligent model, the object features and changes are automatically identified and described, and priority tags are added to the video clips according to the user's personalized monitoring needs to optimize the viewing experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0019] Figure 1 is a hardware environment diagram of a method for determining a query result according to an embodiment of the present application;

[0020] Figure 2 is a flowchart of a method for determining a query result according to an embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of the architecture of a data intelligent push and query system based on intelligent language processing of image data according to an embodiment of the present application;

[0022] Figure 4is a structural block diagram of a device for determining a query result according to an embodiment of the present application;

[0023] Figure 5 is a block diagram of a computer system structure of an electronic device according to an embodiment of the present application;

[0024] Figure 6 An electronic device is provided for implementing a method for determining a query result according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] According to one aspect of an embodiment of the present application, a method for determining a query result is provided. The method for determining a query result is widely used in smart home (Smart Home), smart home, smart home device ecology, smart residential (Intelligence House) ecology and other whole-house intelligent digital control application scenarios. Optionally, in this embodiment, the above-mentioned method for determining a query result can be applied to Figure 1 In the hardware environment composed of the terminal device 102 and the server 104 shown in FIG. Figure 1As shown, the server 104 is connected to the terminal device 102 via a network, and can be used to provide services (such as application services, etc.) for the terminal or a client installed on the terminal. A database can be set on the server or independently of the server to provide data storage services for the server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data computing services for the server 104.

[0028] The network may include but is not limited to at least one of the following: wired network, wireless network. The wired network may include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network, and the wireless network may include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal device 102 may be but is not limited to a PC, a mobile phone, a tablet computer, a smart air conditioner, a smart range hood, a smart refrigerator, a smart oven, a smart stove, a smart washing machine, a smart water heater, a smart washing device, a smart dishwasher, a smart projection device, a smart TV, a smart clothes drying rack, a smart curtain, a smart audio and video, a smart socket, a smart speaker, a smart fresh air device, a smart kitchen and bathroom device, a smart bathroom device, a smart sweeping robot, a smart window cleaning robot, a smart mopping robot, a smart air purification device, a smart steamer, a smart microwave oven, a smart kitchen treasure, a smart purifier, a smart water dispenser, a smart door lock, etc.

[0029] In this embodiment, a method for determining a query result is provided, which can be applied to a computer terminal or an Internet of Things cloud. Figure 2 : is a flowchart of a method for determining a query result according to an embodiment of the present application, the process comprising the following steps:

[0030] S202, obtaining original audio and video data collected by the target device;

[0031] S204, separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data;

[0032] S206, inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes;

[0033] Optionally, in a home monitoring scenario, if the original audio and video data contains the behaviors of both family members and pets. The system first separates the sub-audio and video data sets containing human and pet types through target detection technology. The human subset is sent to the face recognition and behavior analysis sub-model to identify the specific identity and ongoing activities of family members, such as "parents are cooking" or "children are doing homework". The pet subset is sent to the pet recognition and behavior analysis sub-model to identify the type and behavior of pets, such as "cat is sleeping" or "dog is playing". Through this separation processing step, the system can analyze the behaviors of family members and pets simultaneously and in parallel, and generate detailed behavior description text. For example, the system can recognize that "the child comes home after school and is watching TV in the living room, and the pet cat is sleeping next to him." In this way, parents can not only understand the movements of their children in real time, but also grasp the status of their pets, which meets the diversified needs of home monitoring, while also optimizing the data processing process, improving the accuracy of recognition and the response speed of the system.

[0034] S208, adding a to-be-viewed tag to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video segments in the original audio and video data.

[0035] Optionally, in a family monitoring scenario, if parents pay special attention to their children's safety and learning, but don't care much about the behavior of their pets. Through intelligent model analysis, the system identifies multiple events, including "the child comes home after school and starts doing homework", "the pet dog plays in the backyard", and "a stranger passes by when no one is at home at night". Based on the recognition results, the system adds tags to be viewed for the video clips according to the parents' personalized monitoring needs (giving priority to the child's movements). For example, the video clip of "the child comes home after school and starts doing homework" is marked as "high priority", "the pet dog plays in the backyard" is marked as "low priority", and "a stranger passes by when no one is at home at night" may also be marked as "high priority" according to the parents' safety settings. This mechanism ensures that parents can pay attention to the most important monitoring information in a timely manner, and the system also effectively manages the viewing priority of audio and video data, making monitoring more intelligent and efficient, and meeting the personalized needs of users.

[0036] Through the above steps, the original audio and video data collected by the target device is obtained; the original audio and video data is separated and processed to obtain multiple sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; the multiple sub-audio and video data sets are input into the target large model to obtain the target recognition result, wherein the target large model includes multiple sub-models, each sub-model corresponds to an object type; the sub-model is used to identify the object features in the sub-audio and video data set and generate text data for describing the object features and feature changes; according to the monitoring requirements set by the management object and the target recognition result, the audio and video clips in the original audio and video data are added with a to-be-viewed tag, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data. The above technical solution solves the problem of not being able to deeply analyze the video content and accurately identify the feature tags to meet the user's personalized information needs for the monitored object. Furthermore, the audio and video data of multiple types of objects are analyzed by the intelligent model, the object features and changes are automatically identified and described, and priority tags are added to the video clips according to the user's personalized monitoring needs to optimize the viewing experience.

[0037] In an exemplary embodiment, tags to be viewed are added to audio and video clips in the original audio and video data according to monitoring requirements set for the management object and target recognition results, including: parsing the monitoring requirements to obtain characteristics of the object to be monitored; wherein the object characteristics are used to indicate the type characteristics of the target object to be monitored and the specific behavior characteristics of the target object to be monitored; filtering the result content that meets the object characteristics from the target recognition results; filtering out multiple audio and video clips from the original audio and video data based on the correspondence between the result content and the original audio and video data, and adding viewing tags to the multiple audio and video clips respectively, wherein the viewing tags are used to identify the audio and video to be viewed that meet the monitoring requirements and appear in the original audio and video data.

[0038] Optionally, in a home monitoring scenario, if the main monitoring needs of the parents (management objects) are to focus on "children doing homework" and "pet cats eating". The system first parses these monitoring needs and obtains specific object features: "children", "doing homework", "pet cats", and "eating". Next, the system performs real-time analysis on the video captured by the home monitoring camera, and the target recognition results include various activities of family members and the behavior of pets. For example, events such as "children sitting at the desk writing", "children leaving the living room", "pet cats approaching the food bowl", and "pet cats leaving the food bowl" are recognized. The system filters out the result content related to "children doing homework" and "pet cats eating" according to the monitoring needs, such as the two scenes of "children sitting at the desk writing" and "pet cats approaching the food bowl". Based on the filtered result content, the system accurately finds the corresponding video clips from the original audio and video data, and adds a "view tag" to each clip. For example, the video clip containing the content of "children doing homework" is labeled "high priority-children learning", and the clip capturing the scene of "pet cats eating" is labeled "pet-eating". Parents can immediately see notifications with labels such as "High Priority - Child Learning" and "Pet - Eating" on the App. After clicking on the notification, they will be directly redirected to the corresponding video clip playback. In this way, parents can quickly understand their children's learning situation and their pets' eating status without spending time browsing all the videos, ensuring the timeliness and effectiveness of monitoring information while meeting the user's personalized monitoring needs.

[0039] In an exemplary embodiment, before adding viewing tags to multiple audio and video clips respectively, the above method also includes: determining the audio and video correlation between any two clips among the multiple audio and video clips; when the audio and video correlation is greater than or equal to a preset correlation, determining the associated audio and video clips as the same type of events to which the same viewing tag is to be added; when the audio and video correlation is less than the preset correlation, determining to add a different viewing tag for each audio and video clip.

[0040] Optionally, in a home surveillance scenario, several audio and video clips capture a series of activities of a child after returning home from school. Clip A shows "the child enters the house", Clip B captures "the child places a schoolbag in the living room", and Clip C records "the child starts doing homework at the desk". Through analysis, the system determines that the audio and video correlation between Clip A and Clip B is high (for example, the child's continuous actions), and the correlation between Clip C and Clip A and B is also high (because they all occur on the continuous timeline after the child returns home from school), but the correlation between Clip C and Clip B is relatively low (there may be a period of time between placing the schoolbag and doing homework). In this case, the system determines Clip A and B as the same type of events with high correlation, and adds the same viewing label "child returns home-continuous actions" to them. As for Clip C, although it is associated with A and B, based on the preset correlation threshold, the system believes that Clip C describes a new behavior stage, so a different viewing label "child starts studying" is added to Clip C. When parents check on the app, they will see a notification with the label "Child Goes Home - Continuous Actions", showing the process of the child going home and putting down the schoolbag, and a separate notification with the label "Child Starts Studying", clearly informing the child that he has started doing homework. This classification method ensures that parents can clearly understand the child's movement trajectory, while avoiding unnecessary repeated notifications and improving the management and use efficiency of monitoring information.

[0041] In an exemplary embodiment, after adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition results, the above method also includes: determining the target attention features corresponding to the set monitoring requirements; labeling the target attention features to obtain multiple attention tags; matching the multiple attention tags with all the viewing tags that have been added; and determining a text message to be pushed to the management object of the target device based on the matching results, wherein the text message is used to carry a description text of the audio and video content that meets the set monitoring requirements.

[0042] Optionally, in a smart home monitoring system, the monitoring requirements set by parents (management objects) are to pay attention to "children's activities after school" and "abnormal behavior of pet cats". The system first converts these requirements into specific feature tags, such as "school time", "activity type", "pet abnormality", etc. Subsequently, the system performs target recognition and analysis on the captured audio and video clips, and adds viewing tags to each clip, for example, "Clip X: Child returns home", "Clip Y: Child plays in the living room", "Clip Z: Pet cat wanders at the door", etc. Then, the system matches the monitoring requirement attention tags with all added viewing tags. In this embodiment, "school time" and "activity type" are highly matched with the viewing tags of clips X and Y, while "pet abnormality" matches the viewing tag of clip Z. Based on the above matching results, the system generates and pushes two text messages: one for the child's after-school activities, with the message description "Your child has arrived home safely after school and is playing in the living room." The other for the abnormal behavior of the pet cat, with the message description "Please note that the pet cat is wandering at the door and may be behaving abnormally." Parents receive these two text messages on the mobile app and immediately understand the child's status after school and the abnormal behavior of the pet cat. There is no need to manually check each segment, which greatly saves time and ensures the timeliness and relevance of monitoring information. This intelligent push mechanism based on tag matching makes the home monitoring system more in line with the personalized needs of users and provides efficient and convenient information services.

[0043] In an exemplary embodiment, after adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements and target recognition results set by the management object, the above method also includes: when receiving an audio and video viewing request input by other objects, determining the target tag feature corresponding to the current audio and video viewing request; using the target tag feature to search in the target database, and synchronizing the search results to the target terminal corresponding to the other object for display, wherein the target database is used to store the original audio and video data to which the tags to be viewed have been added.

[0044] Optionally, in a smart home monitoring system, another family member (such as a family nanny) initiates an audio and video viewing request through his mobile phone App (target terminal) to view the content related to "children doing homework". The system first analyzes the request and determines that its corresponding target tag features are "children" and "doing homework". Subsequently, the system uses these features to search in the target database, which stores all the original audio and video data to which the tags to be viewed have been added. In this embodiment, the database contains multiple audio and video clips, such as "children go home from school", "children play in the living room", "children start doing homework", etc., where the viewing tag of the "children start doing homework" clip matches the target tag feature of the current request. The system synchronizes the retrieval results of the clips related to "children doing homework" to the nanny's mobile phone App for display. The displayed content may include a summary description of the video clip, a timestamp, and a brief text description, such as "[2023-04-05 16:30] The child has sat at the desk and started doing homework." The nanny can immediately view the video clip on the App to understand the child's latest status without having to view a large amount of video data unrelated to the request. Through this process, the smart home monitoring system can efficiently and accurately respond to family members' audio and video viewing requests, provide instant and personalized content display, meet the specific needs of different members for monitoring information, and at the same time optimize the use of system resources to ensure smooth system operation and responsiveness.

[0045] In an exemplary embodiment, before searching in the target database using the target tag feature, the method further includes: determining a first search permission for other objects corresponding to the audio and video viewing request; allowing other objects to perform real-time viewing of audio and video data when the first search permission is the same as a second search permission set for the management object; prohibiting other objects from performing real-time viewing of audio and video data when the first search permission is different from the second search permission set for the management object, and notifying the management object of the request data of other objects.

[0046] Optionally, in a smart home monitoring system, the father in the family is the management object. He sets the second search permission in the system, allowing the mother to view the audio and video data of all cameras in real time, but restricting the neighbor to view the video of the doorbell camera only after the father's confirmation. One day, the mother initiated a request to view the audio and video data of the living room camera through the mobile phone App. The system first determines the mother's first search permission and finds that she is authorized to view the data of all cameras in real time, which matches the second search permission set by the father. Therefore, the system allows the mother to view the audio and video data of the living room camera in real time without additional confirmation. However, the neighbor tries to view the real-time video of the doorbell camera through his mobile phone App on the same day. The system determines the neighbor's first search permission and finds that the neighbor does not have the default real-time viewing permission, which does not match the father's second search permission. Therefore, the system prohibits the neighbor from viewing the data of the doorbell camera in real time, and notifies the father of the neighbor's request data, waiting for the father's further confirmation or authorization. Through the meticulous control and comparison of permissions, the smart home monitoring system ensures the security, transparency and trust between family members of the monitoring information, while also providing an instant channel for information acquisition in emergency situations, enhancing the effectiveness and flexibility of home monitoring.

[0047] In an exemplary embodiment, before adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements and target recognition results set by the management object, the above method also includes: counting the changes in historical monitoring requirements; dynamically adjusting the generation rules of the tags to be viewed according to the changes in requirements, wherein the generation rules include at least: the type of tags to be viewed, the description of the tags to be viewed, and the identification criteria corresponding to the tags to be viewed; and updating the tag content of the tags to be viewed based on the generation rules.

[0048] Optionally, in a smart home monitoring system, the system records the changes in the monitoring needs of family members in the past month. Through data analysis, it is found that family members pay more attention to the "family gathering" scene on weekends, while they pay more attention to "children coming home from school" and "pet cat behavior" on weekdays. Based on this change in demand, the system dynamically adjusts the generation rules of the tags to be viewed, adding recognition criteria for the "family gathering" tag for weekends, such as identifying scenes such as the gathering of family members and food on the table. At the same time, the generation rules for weekdays focus on the recognition of "children coming home from school" and "pet cat behavior", including a more detailed analysis of the child's facial expression, the location of the school bag, and the activity area and behavior pattern of the pet cat. With the implementation of the new rules, the system began to generate the "family gathering" tag in weekend video surveillance, and more "children coming home from school" and "pet cat behavior" related tags in weekday surveillance. When family members view historical monitoring data, they can intuitively see the labels automatically generated by the system based on their changing needs. For example, on weekday evenings, mothers can directly browse video clips with labels such as "children coming home from school" and "pet cat behavior" without having to view the "family gathering" scenes of the entire weekend. By dynamically adjusting the generation rules according to changes in needs, the system not only improves the personalization and relevance of monitoring information, but also improves the efficiency of family members in obtaining monitoring information, ensuring the efficient operation of the intelligent monitoring system and the satisfactory experience of family members. This flexible adjustment mechanism enables the intelligent monitoring system to better adapt to the dynamic changes in family life and provide more accurate and intelligent services.

[0049] In order to better understand the process of determining the query results, the implementation method of calling the above data resources is described below in conjunction with an optional embodiment, but it is not used to limit the technical solution of the embodiment of the present application.

[0050] In the related art, the data generated by video surveillance are all data presentations of a single subject, and no more detailed and concrete analysis and annotation are performed on the subject appearing in the data. In order to solve the above problems, the optional embodiment of the present application proposes a data intelligent push and query method based on intelligent language processing of image data. The method starts with the video data content captured and reported by the audio and video equipment, separates the video content with human shape, human face and pet, and performs data tagging. Different data are sent to the intelligent language model according to type for feature content recognition, and the accuracy of data recognition is guaranteed. According to the personalized attention features set by the user, the data characterized and annotated by the intelligent language model is filtered for relevance, and the data that meets the personalized attention features of the user is summarized into related text message content, which is pushed to the App of the related user in real time. At the same time, in the App, the user can initiate a historical data query based on conditions such as time, specific feature tags or monitoring areas. After receiving the query request, the system quickly retrieves the matching information in the database, and presents the relevant video clips and their feature descriptions to the user, so that the user can view the action trajectory and real-time behavior of the target object.

[0051] Optional, Figure 3 This is a schematic diagram of the architecture of a data intelligent push and query system based on intelligent language processing of image data according to an embodiment of the present application. The key components involved in the architecture include alarm message reporting, Internet of Things (IoT) message platform, message subscription, local audio and video recognition, device (client), application (Application, App), user tag data query, alarm message push, data storage, audio and video services, server, scheduled tasks, video structured tags, and audio and video data storage services. Specifically, it can be executed according to the following steps:

[0052] Step 1: Device alarm and video upload description. When a smart camera or other audio and video device detects a preset abnormal situation, such as the appearance of a human figure, face, or pet, or a change in the image in the monitoring area, the device will trigger an alarm and start recording a 7-second video. Subsequently, this video and alarm information are uploaded to the IoT (Internet of Things) data platform through the data channel. As the source of information, the device ensures the timely generation and transmission of monitoring data, laying the foundation for subsequent analysis and processing.

[0053] Step 2: IoT message platform receiving and transferring description. The IoT message platform receives the video data and alarm messages uploaded by the device. The platform passes this data to the audio and video device subscription service, that is, the audio and video server, through the "message subscription" mechanism. Apps and servers subscribe to alarm messages of specific devices or events through the IoT message platform, and can receive relevant alarm information in the first time for subsequent processing. The IoT message platform plays the role of a data transfer station, efficiently transmitting the data of the front-end device to the back-end processing system.

[0054] Step 3: Intelligent language model analysis description. After receiving the video data, the audio and video server sends the data to the intelligent language model service for feature label recognition based on the type of video monitoring (such as abnormal monitoring area, presence of people, pets, face recognition, etc.). The model analyzes the visual information, audio information, and contextual information in the video to identify character features, movements, behaviors, and pet dynamics. Through the in-depth analysis of the intelligent language model, detailed behavior and feature labels are added to the video data, enhancing the descriptiveness and application value of the data.

[0055] Optionally, the intelligent language model can identify the characteristic content of different types of data processing and ensure the accuracy of data identification through the following technologies. The details are as follows:

[0056] (1) Multimodal information fusion. This includes at least but is not limited to the following three types of information: visual information, audio information, and contextual information. Visual information includes: face, posture, clothing, behavior, etc. Audio information includes at least: voice, voiceprint, etc. Contextual information includes at least: time, location, scene, etc.

[0057] (2) Balance between real-time and accuracy. Specifically, real-time is reflected in the fact that the monitoring system requires that the identification and behavior analysis of people be performed in real time. Accuracy is to ensure the accuracy of push messages, which requires accurate identification of people's characteristics, movements and behaviors.

[0058] (3) Robustness in complex scenes. Specific scenes include: illumination changes, occlusion problems, motion blur, and background interference. Among them, illumination changes will affect image quality under different lighting conditions; occlusion problems refer to partial or complete occlusion of the human body; motion blur is caused by the rapid movement of the human body, resulting in blurred images; and background interference is caused by the fact that complex backgrounds easily interfere with target detection.

[0059] Optionally, the software or hardware on the device can perform preliminary local recognition of the collected raw audio and video data, for example, identifying different types of objects such as faces, vehicles, sounds, etc. in the picture. Local recognition can respond quickly, but may be limited by the computing power of the device. After local recognition, the device stores the raw audio and video data and the recognition results for subsequent retrieval or analysis. Data storage can be local storage on the device or a remote data storage service connected via a network.

[0060] Optionally, when the results of the characterization tag recognition or the stored data of the intelligent language model service contain urgent or important events, the server will generate an alarm message and push these messages to related apps or devices through the network to ensure that users can learn about the events they are concerned about in a timely manner.

[0061] Step 4: Tag data transmission and personalized screening description: After the intelligent language model is processed, the feature tags and video description data are transmitted back to the audio and video server. The audio and video server performs correlation screening and filtering on the data according to the personalized attention characteristics of the user, such as parents paying attention to children's behavior and pet lovers paying attention to pet dynamics, to form message content with specific tags. This ensures the personalization and accuracy of the message content, and users can receive real-time updated information that matches their needs.

[0062] Step 5: Message push and data storage description: The audio and video server pushes the filtered and screened message content to the App. Users can click on these messages to jump to the corresponding monitoring data for viewing and understanding. At the same time, the audio and video server saves the labeled data in the database for users to retrieve historical data on the App. This achieves instant information push and traceability of historical data, improving user experience and monitoring data management efficiency.

[0063] Step 6: Description of scheduled tasks and batch data processing. The system sets a scheduled task, which is executed at 0:00 am every day to process all messages of the day and push them in batches. The audio and video server interacts with the video structured tag service to tag all video data of the day, and then pushes these tagged data to users through the App to complete batch data processing and user notification. The regular combing and data update of the system are ensured, and users can receive summary information regularly to facilitate long-term monitoring and behavior analysis. In addition, scheduled tasks can also include regular backup of audio and video data, clearing expired data, and other operations.

[0064] It should be noted that when uploading data, it is necessary to ensure that the recorded video and alarm information are uploaded to the IoT data platform using a secure data transmission protocol, such as HTTPS or TLS, to prevent the data from being intercepted or tampered with during transmission. At the same time, data storage costs and network transmission efficiency need to be considered. If the video quality is poor, it may affect subsequent feature recognition. In addition, the subscription service needs to be able to accurately select the corresponding intelligent language model for feature label recognition based on the type of video data to improve the pertinence and accuracy of recognition. Business services need to accurately understand the personalized needs of users, such as parents paying attention to their children's behavior and pet owners paying attention to pet dynamics, to ensure that the pushed messages are highly matched with the user's focus. The generated text messages should be concise, clear, easy to understand, avoid technical terms or complex descriptions, and ensure that users can quickly obtain the key information they need. When storing feature data in the database, measures need to be taken to protect user privacy, such as data desensitization and anonymization, and ensure compliance with data protection regulations.

[0065] Through the above scheme, the deep learning algorithm is used to realize the intelligent recognition and in-depth analysis of the behavioral characteristics of people and pets in the video clips, and the multimodal information fusion method is used to improve the accuracy and robustness of the recognition. Specifically, when the audio and video equipment detects specific abnormal activities, it automatically records the video and uploads it to the IoT data platform, and then the data is sent to the intelligent language model for feature label recognition. The business service screens the feature data according to the user's personalized settings, summarizes the information that meets the set conditions into text messages, and accurately pushes it to the user through the App, while storing the data securely. The above technical solution significantly improves the efficiency and accuracy of users in obtaining monitoring information, reduces the push of invalid information, enhances the adaptability of the system in complex monitoring environments, meets the diverse needs of family members for monitoring content, and ensures the security and privacy of user data through encrypted transmission and permission control.

[0066] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of each embodiment of the present application.

[0067] Figure 4 is a structural block diagram of a device for determining a query result according to an embodiment of the present application; Figure 4As shown, including:

[0068] The acquisition module is used to obtain the original audio and video data collected by the target device;

[0069] A processing module 44 is used to separate and process the original audio and video data to obtain multiple sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data;

[0070] An input module 46 is used to input the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes;

[0071] The adding module 48 is used to add a to-be-viewed tag to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video segments in the original audio and video data.

[0072] Through the above device, the original audio and video data collected by the target device is obtained; the original audio and video data is separated and processed to obtain multiple sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; the multiple sub-audio and video data sets are input into the target large model to obtain the target recognition result, wherein the target large model includes multiple sub-models, each sub-model corresponds to an object type; the sub-model is used to identify the object features in the sub-audio and video data set and generate text data for describing the object features and feature changes; according to the monitoring requirements set by the management object and the target recognition result, the audio and video clips in the original audio and video data are added with a to-be-viewed tag, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data. The above technical solution solves the problem of being unable to deeply analyze the video content and accurately identify the feature tags to meet the user's personalized information needs for the monitored object. Furthermore, the audio and video data of multiple types of objects are analyzed by the intelligent model, the object features and changes are automatically identified and described, and priority tags are added to the video clips according to the user's personalized monitoring needs to optimize the viewing experience.

[0073] In an exemplary embodiment, the above-mentioned adding module is also used to parse the monitoring requirements and obtain the characteristics of the object to be monitored; wherein the object characteristics are used to indicate the type characteristics of the target object to be monitored and the specific behavioral characteristics of the target object to be monitored; the result content that meets the object characteristics is filtered out from the target recognition results; based on the correspondence between the result content and the original audio and video data, multiple audio and video clips are filtered out from the original audio and video data, and viewing tags are added to the multiple audio and video clips respectively, wherein the viewing tags are used to identify the audio and video to be viewed that meet the monitoring requirements and appear in the original audio and video data.

[0074] In an exemplary embodiment, the above-mentioned adding module also includes: a first determination unit, used to determine the audio and video correlation between any two clips among multiple audio and video clips; when the audio and video correlation is greater than or equal to a preset correlation, the associated audio and video clips are determined as the same type of events to which the same viewing tag is to be added; when the audio and video correlation is less than a preset correlation, it is determined to add a different viewing tag for each audio and video clip.

[0075] In an exemplary embodiment, the above-mentioned device also includes: a first determination module, which is used to add a to-be-viewed label to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, and then determine the target attention feature corresponding to the set monitoring requirement; label the target attention feature to obtain multiple attention labels; match the multiple attention labels with all the viewing labels that have been added; determine the text message to be pushed to the management object of the target device according to the matching result, wherein the text message is used to carry the audio and video content description text that meets the set monitoring requirements.

[0076] In an exemplary embodiment, the above-mentioned device also includes: a second determination module, which is used to add a to-be-viewed label to the audio and video clips in the original audio and video data according to the monitoring requirements and target recognition results set by the management object, and then, when receiving an audio and video viewing request input by other objects, determine the target label feature corresponding to the current audio and video viewing request; use the target label feature to search in the target database, and synchronize the search results to the target terminal corresponding to the other object for display, wherein the target database is used to store the original audio and video data to which the to-be-viewed label has been added.

[0077] In an exemplary embodiment, the above-mentioned second determination module also includes: a second determination unit, which is used to determine the first search permission of other objects corresponding to the audio and video viewing request before using the target tag feature to search in the target database; when the first search permission is the same as the second search permission set for the management object, other objects are allowed to perform real-time viewing of audio and video data; when the first search permission is different from the second search permission set for the management object, other objects are prohibited from performing real-time viewing of audio and video data, and the request data of other objects are notified to the management object.

[0078] In an exemplary embodiment, the above-mentioned device also includes: a statistical module, which is used to count the changes in historical monitoring needs before adding tags to be viewed to the audio and video clips in the original audio and video data according to the monitoring needs set by the management object and the target recognition results; dynamically adjust the generation rules of the tags to be viewed according to the changes in needs, wherein the generation rules include at least: the type of tags to be viewed, the description of the tags to be viewed and the identification criteria corresponding to the tags to be viewed; and update the tag content of the tags to be viewed based on the generation rules.

[0079] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.

[0080] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0081] An embodiment of the present application further provides a storage medium, which includes a stored program, wherein the program executes any of the above methods when it is run.

[0082] Optionally, in this embodiment, the storage medium may be configured to store program codes for executing the following steps:

[0083] S1, obtaining the original audio and video data collected by the target device;

[0084] S2, separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data;

[0085] S3, inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes;

[0086] S4, adding a to-be-viewed tag to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data.

[0087] An embodiment of the present application further provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0088] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0089] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:

[0090] S1, obtaining the original audio and video data collected by the target device;

[0091] S2, separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data;

[0092] S3, inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes;

[0093] S4, adding a to-be-viewed tag to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data.

[0094] Optionally, in this embodiment, the above-mentioned storage medium may include but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and other media that can store program codes.

[0095] Figure 5 The computer system structure block diagram for implementing the electronic device of the embodiment of the present application is schematically shown. It should be noted that: Figure 5 The computer system 500 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application. Figure 5 As shown, the computer system 500 includes a central processing unit 501 (CPU), which can perform various appropriate actions and processes according to the program stored in the read-only memory 502 (ROM) or the program loaded from the storage part 508 to the random access memory 503 (RAM). Various programs and data required for system operation are also stored in the random access memory 503. The central processing unit 501, the read-only memory 502 and the random access memory 503 are connected to each other through a bus 504. The input / output interface 505 (Input / Output interface, i.e., I / O interface) is also connected to the bus 504.

[0096] The following components are connected to the input / output interface 505: an input section 505 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a local area network card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the input / output interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read therefrom is installed into the storage section 508 as needed.

[0097] In particular, according to an embodiment of the present application, the process described in each method flow chart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer readable medium, and the computer program contains a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processor 501, various functions defined in the system of the present application are executed.

[0098] According to another aspect of the embodiment of the present application, an electronic device for implementing the above-mentioned data resource call is also provided. Figure 6 As shown, the electronic device includes a memory 602 and a processor 604. The memory 602 stores a computer program, and the processor 604 is configured to execute the steps in any of the above method embodiments through the computer program.

[0099] Optionally, in this embodiment, the electronic device may be located in at least one of a plurality of network devices of a computer network. A person skilled in the art may understand that Figure 6 The structure shown is for illustration only, and the electronic device may also be a device including the above-mentioned flash memory. Figure 6 The structure of the electronic device is not limited. Figure 6 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 6 Different configurations shown.

[0100] Among them, the memory 602 can be used to store software programs and modules, such as the call of data resources in the embodiment of the present application and the program instructions / modules corresponding to the device. The processor 604 executes various functional applications and data processing by running the software programs and modules stored in the memory 602, that is, realizing the call of the above-mentioned data resources. The memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 602 may further include a memory remotely arranged relative to the processor 604, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 602 can specifically be, but is not limited to, used for information such as logs containing modeling data. As an example, such as Figure 6As shown, the memory 602 may include but is not limited to the modules in the device for calling the data resource. In addition, it may also include but is not limited to other module units in the device for calling the data resource, which will not be described in detail in this example.

[0101] Optionally, the transmission device 606 is used to receive or send data via a network. Specific examples of the above-mentioned network may include wired networks and wireless networks. In one example, the transmission device 606 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 606 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet wirelessly. In addition, the above-mentioned electronic device also includes: a display 608; and a connection bus 610, which is used to connect the various module components in the above-mentioned electronic device.

[0102] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any one of the above method embodiments are implemented.

[0103] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in any of the above method embodiments are implemented.

[0104] The embodiments of the present application also provide a computer program, which includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of any one of the above method embodiments.

[0105] Optionally, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation modes, and this embodiment will not be described in detail here.

[0106] Obviously, those skilled in the art should understand that the above modules or steps of the present application can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order from that herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.

[0107] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for determining a query result, characterized in that: include: Obtain the original audio and video data collected by the target device; Separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; Inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used to identify object features in the sub-audio and video data sets and generate text data for describing the object features and feature changes; According to the monitoring requirements set by the management object and the target recognition result, a to-be-viewed tag is added to the audio and video clips in the original audio and video data, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data.

2. The method for determining the query result according to claim 1, characterized in that: Adding a tag to be viewed to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result includes: Parsing the monitoring requirements to obtain the characteristics of the object to be monitored; wherein the object characteristics are used to indicate the type characteristics of the target object to be monitored and the specific behavior characteristics of the target object to be monitored; Filtering the result content that meets the characteristics of the object from the target recognition results; Based on the correspondence between the result content and the original audio and video data, multiple audio and video clips are screened out from the original audio and video data, and viewing tags are added to the multiple audio and video clips respectively, wherein the viewing tags are used to identify the audio and video to be viewed that meet the monitoring requirements and appear in the original audio and video data.

3. The method for determining the query result according to claim 2, characterized in that: Before adding viewing tags to the multiple audio and video clips respectively, the method further includes: Determine the audio and video correlation between any two of the multiple audio and video clips; When the audio and video correlation degree is greater than or equal to a preset correlation degree, the associated audio and video clips are determined as the same type of events to be added with the same viewing tag; When the audio and video correlation is less than a preset correlation, it is determined to add a different viewing tag for each audio and video segment.

4. The method for determining a query result according to claim 1, characterized in that: After adding a to-be-viewed tag to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, the method further includes: Determine the target focus feature corresponding to the set monitoring requirement; Labeling the target attention features to obtain multiple attention labels; Matching the multiple attention tags with all the viewing tags that have been added; A text message to be pushed to the management object of the target device is determined according to the matching result, wherein the text message is used to carry a description text of the audio and video content that meets the set monitoring requirements.

5. The method for determining a query result according to claim 1, characterized in that: After adding a to-be-viewed tag to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, the method further includes: When receiving an audio or video viewing request input by another object, determining a target tag feature corresponding to the current audio or video viewing request; The target tag feature is used to search in a target database, and the search result is synchronized to the target terminal corresponding to the other object for display, wherein the target database is used to store the original audio and video data to which the tag to be viewed has been added.

6. The method for determining the query result according to claim 5, characterized in that: Before searching in a target database using the target tag feature, the method further includes: Determine the first search permission of other objects corresponding to the audio and video viewing request; When the first search permission is the same as the second search permission set for the management object, the other object is allowed to view the audio and video data in real time; In the case where the first search authority is different from the second search authority set for the management object, the other objects are prohibited from viewing the audio and video data in real time, and the management object is notified of the request data of the other objects.

7. The method for determining a query result according to claim 1, characterized in that: Before adding a to-be-viewed tag to the audio and video segments in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, the method further includes: Statistical history monitors demand changes; Dynamically adjusting the generation rules of the tags to be viewed according to the demand changes, wherein the generation rules at least include: the type of the tags to be viewed, the description of the tags to be viewed, and the identification standard corresponding to the tags to be viewed; The tag content of the tag to be viewed is updated based on the generation rule.

8. A device for determining a query result, characterized in that: include: An acquisition module is used to obtain the original audio and video data collected by the target device; A processing module, used for separating and processing the original audio and video data to obtain a plurality of sub-audio and video data sets, wherein each sub-audio and video data set corresponds to an object type, and there are at least two object types in the original audio and video data; An input module, used for inputting the plurality of sub-audio and video data sets into a target large model to obtain a target recognition result, wherein the target large model includes a plurality of sub-models, each sub-model corresponding to an object type; the sub-model is used for identifying object features in the sub-audio and video data sets and generating text data for describing the object features and feature changes; An adding module is used to add a to-be-viewed tag to the audio and video clips in the original audio and video data according to the monitoring requirements set by the management object and the target recognition result, wherein the to-be-viewed tag is used to manage the priority of viewing different audio and video clips in the original audio and video data.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 7 when executed.

10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 7 through the computer program.