Scene understanding method and system and electronic equipment

By continuously acquiring multi-dimensional scene information and conducting multimodal analysis, we obtain comprehensive scene understanding information, solve the problem of single analysis results of existing smart hardware devices, and achieve deep interaction with users and effective monitoring of target object learning.

CN120656233APending Publication Date: 2025-09-16BEIJING XUEDIRUANJIAN DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510719676.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing smart hardware devices can only process image information at a specific moment, the analysis results are single, and there is a lack of further interaction with users, which cannot meet the diverse practical application needs.

Method used

By continuously acquiring multi-dimensional scene information, including image information, voice information and text information, and using preset multimodal analysis strategies to perform scene analysis, multi-dimensional scene understanding information, such as attribute information, relationship information, behavior information and emotional information, is obtained, and interaction with users is carried out based on this information.

Benefits of technology

It achieves a more comprehensive and detailed understanding of the scene, supports better interaction with users, can better monitor the learning status of the target objects, assist their learning, and meet diverse practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656233A_ABST
    Figure CN120656233A_ABST
Patent Text Reader

Abstract

The invention discloses a scene understanding method and system and electronic equipment, and relates to the field of artificial intelligence, and the method comprises the steps: continuously obtaining multi-dimensional scene information, carrying out the scene analysis of image information, voice information and text information according to a preset multi-modal analysis strategy, so as to obtain multi-dimensional scene understanding information, and carrying out the analysis of the multi-dimensional scene understanding information. The multi-dimensional scene understanding information comprises attribute information representing setting conditions of objects in the monitoring scene, relation information representing association conditions among the objects, behavior information of the target monitoring object and emotion information of the target monitoring object; and interacting with the user based on the multi-dimensional scene understanding information. According to the scheme, the multi-dimensional scene information can be continuously acquired so as to correspondingly determine the multi-dimensional scene understanding information, and the continuous acquisition is beneficial to monitoring the target monitoring object from a long-time dimension; the analysis result relates to four dimensions, so that scene understanding information is more comprehensive, subsequent better interaction with a user is facilitated, and better monitoring of the learning condition of the target monitoring object is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a scene understanding method, system, and electronic device. Background Art

[0002] With the continuous vigorous development of artificial intelligence technology, smart hardware devices obtained by hardware-based educational products have become an important source of students' independent learning. However, current smart hardware devices can only process image information of specific objects at specific moments, that is, conduct behavioral analysis to determine the behavior of specific objects at the current moment. The analysis results are relatively simple and less comprehensive, lack further interaction with users, and lack long-term monitoring of specific objects from a time dimension, which cannot meet the increasingly diverse practical application needs. Summary of the Invention

[0003] In view of this, the present invention provides a scene understanding method, system and electronic device, which provide more comprehensive scene understanding information, facilitate better subsequent interaction with users, and facilitate better monitoring of the learning status of the target monitored object.

[0004] To solve the above technical problems, this application provides a scene understanding method, including:

[0005] Continuously acquiring multi-dimensional scene information, wherein the multi-dimensional scene information includes image information, voice information, and text information;

[0006] Performing scene analysis on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, wherein the multidimensional scene understanding information includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between the objects, behavior information of the target monitoring object, and emotional information of the target monitoring object;

[0007] Interact with the user based on the multi-dimensional scene understanding information.

[0008] Furthermore, interacting with the user based on the multi-dimensional scene understanding information includes:

[0009] When the behavior recorded in the behavior information continues to appear within a first preset time period, determining that the behavior is a learning habit of the target monitored object;

[0010] Determining whether the learning habit is a bad learning habit based on a preset habit determination benchmark;

[0011] If so, the result that the target monitoring object has the bad learning habit is stored, and / or the prompt module is controlled to prompt the target monitoring object that the bad learning habit exists.

[0012] Furthermore, after obtaining multi-dimensional scene understanding information, it also includes:

[0013] Store the behavior information and / or emotion information of the target monitoring object at each moment.

[0014] Furthermore, interacting with the user based on the multi-dimensional scene understanding information includes:

[0015] According to the preset dangerous behavior judgment criteria, determine whether there is any dangerous behavior in the behavior recorded in the current behavior information;

[0016] If so, the control prompt module prompts the target monitored object to stop performing the dangerous behavior.

[0017] Furthermore, interacting with the user based on the multi-dimensional scene understanding information includes:

[0018] According to a preset long-time behavior reminder benchmark, determining whether there is a target behavior in the behavior recorded in the behavior information that has lasted for a second preset time period;

[0019] If so, the control prompt module prompts the target monitoring object that the time duration for performing the target behavior has reached the second preset time duration.

[0020] Furthermore, scene analysis is performed on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multi-dimensional scene understanding information, including:

[0021] Processing the image information according to N preset image processing algorithms respectively to obtain N basic image information, where N is an integer greater than 1;

[0022] Preprocessing the voice information to obtain corresponding basic voice information;

[0023] Preprocessing the text information to obtain corresponding basic text information;

[0024] The basic image information, the basic voice information and the basic text information are processed according to a preset multimodal large model to obtain multi-dimensional scene understanding information.

[0025] Furthermore, the basic image information, the basic voice information, and the basic text information are processed according to a preset multimodal large model to obtain multi-dimensional scene understanding information, including:

[0026] Input N basic image information as input to a pre-trained image encoding model to extract image features;

[0027] Processing the image features using an image-text alignment module to obtain the image features in a text-comprehensible form;

[0028] Inputting the basic speech information as an input item into a pre-trained speech coding model to extract speech features;

[0029] Processing the speech features using a speech-to-text alignment module to obtain the speech features in a text-comprehensible form;

[0030] Inputting the basic text information as input into a pre-trained text encoding model to extract text features;

[0031] The text features, the image features of the text-understandable form, and the speech features are input as input items to a pre-trained decoding model to obtain multi-dimensional scene understanding information.

[0032] Furthermore, N kinds of basic image information are input as input items to a pre-trained image encoding model to extract image features, including:

[0033] Using the first network to extract global features from N kinds of basic image information to obtain global features;

[0034] Using a second network, local feature extraction is performed on the image sub-regions obtained after each of the N types of basic image information is fragmented to obtain local features;

[0035] The global features and the local features are fused using a preset feature fusion strategy to obtain image features.

[0036] Furthermore, the text features, the image features of the text-understandable form, and the speech features are input as input items to a pre-trained decoding model to obtain multi-dimensional scene understanding information, including:

[0037] Inputting the text features, the image features of the understandable form of the text, and the speech features as input items into a pre-trained text decoding model to obtain attribute information, relationship information, first behavior information, and first emotion information;

[0038] The text features and the speech features of the text-understandable form are input as input items to a pre-trained speech decoding model to obtain second behavior information and second emotion information.

[0039] Furthermore, after obtaining the text features and the image features of the comprehensible form of the text, the method further includes:

[0040] The text features and the image features in the understandable form of the text are input as input items to a pre-trained image decoding module to obtain an image segmentation map for identifying each object in the image information.

[0041] Furthermore, after obtaining the image segmentation map for identifying each object in the image information, the method further includes:

[0042] Annotating each identified object in the image information according to the image segmentation map, and outputting the annotation result through a display module;

[0043] When a search flag signal indicating that any object given in the annotation result has been parsed is received, a network search is performed in response to the search flag signal to obtain and output a parsing result corresponding to the object.

[0044] To solve the above technical problems, the present invention further provides a scene understanding system, comprising:

[0045] A data acquisition unit, configured to continuously acquire multi-dimensional scene information, wherein the multi-dimensional scene information includes image information, voice information, and text information;

[0046] a multimodal processing unit, configured to perform scene analysis on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, wherein the multidimensional scene understanding information includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between the objects, behavioral information of the target monitoring object, and emotional information of the target monitoring object;

[0047] A comprehensive analysis and interaction unit is used to interact with the user based on the multi-dimensional scene understanding information.

[0048] To solve the above technical problems, the present invention further provides an electronic device, comprising:

[0049] memory for storing computer programs;

[0050] A processor is configured to implement the steps of the scene understanding method as described above when executing the computer program.

[0051] The present application provides a scene understanding method, system and electronic device, which continuously acquires multi-dimensional scene information, including image information, voice information and text information, and performs scene analysis on the image information, voice information and text information according to a preset multimodal analysis strategy to obtain multi-dimensional scene understanding information, which includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between each object, behavior information of the target monitoring object and emotional information of the target monitoring object; interacting with the user based on the multi-dimensional scene understanding information. It can be seen that the scheme can continuously acquire multi-dimensional scene information in order to determine the corresponding multi-dimensional scene understanding information. The acquired scene information is more diverse, and the continuous acquisition is conducive to the monitoring of the target monitoring object from a long-term dimension; the analysis results involve four dimensions, making the scene understanding information more comprehensive, which is conducive to better interaction with the user in the future, better monitoring of the learning situation of the target monitoring object, assisting its learning, and facilitating practical application.

[0052] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0054] Figure 1 A flowchart of a scene understanding method provided by the present invention;

[0055] Figure 2 A schematic diagram of the structure of a scene understanding system provided by the present invention;

[0056] Figure 3 This is a structural diagram of an electronic device provided by the present invention. DETAILED DESCRIPTION

[0057] The core of the present invention is to provide a scene understanding method, system and electronic device, which provide more comprehensive scene understanding information, facilitate better interaction with users and better monitoring of the learning status of the target monitoring object.

[0058] The following will be combined with the accompanying drawings in the embodiments of this application to clearly describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field are within the scope of protection of this application.

[0059] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.

[0060] Please refer to Figure 1 , Figure 1 A flowchart of a scene understanding method provided by the present invention;

[0061] The scene understanding method includes:

[0062] S11: Continuously acquiring multi-dimensional scene information, where the multi-dimensional scene information includes image information, voice information, and text information;

[0063] Specifically, the scene understanding method can be applied to electronic devices, which include but are not limited to various computer devices, or intelligent hardware devices related to teaching, such as learning machines, etc., which are not particularly limited here; the electronic device is provided with an image acquisition module, such as a camera, and the image information here can specifically be an image within the monitorable range captured by the image acquisition module, that is, an image of the target monitoring object and the monitoring scene in which the target monitoring object is located. The target monitoring object here can be a student; the electronic device is also provided with a voice acquisition module, such as a microphone, and the voice information here can specifically be audio information input by the user through voice, and / or audio information corresponding to the surrounding monitoring scene, and / or audio information generated when the target monitoring object performs a certain behavior, etc. The user here can be the target monitoring object and / or the guardian of the target monitoring object, such as a student and a parent; the electronic device is also provided with a display module, such as a display screen, which can display a human-computer interaction interface. The text information here can be text information actively input by the user through the human-computer interaction interface. Of course, the text information can also be text information obtained through the communication module on the electronic device and sent by the upper electronic device. The specific content of the information is not particularly limited here and can be flexibly obtained according to actual needs.

[0064] More specifically, step S11 may be to continuously acquire multi-dimensional scene information according to a preset sampling period, and the specific value of the preset sampling period may be flexibly set according to actual needs.

[0065] S12: performing scene analysis on the image information, voice information, and text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, which includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between objects, behavioral information of the target monitoring object, and emotional information of the target monitoring object;

[0066] S13: Interact with users based on multi-dimensional scene understanding information.

[0067] It should be noted that, taking the learning scenario as an example, the attribute information in the multi-dimensional scene understanding information here can be the basic information of the objects included in the monitoring scene (such as category, color, number of settings, etc.), and / or the setting position of the object in the image. For example, the attribute information may include: apple, color is red, number of settings is three, and it is located in the upper left corner of the desktop; relationship information includes: apple is on the desktop; behavior information: the target monitoring object, such as a student, is eating an apple; emotion information: the target monitoring object's emotion is happy.

[0068] In addition, the specific content of the attribute information, relationship information, behavior information and emotional information here may involve multiple objects. Taking the learning scenario as an example, the attribute information may include: the settings of objects such as books, pens, water cups, drinks, erasers, etc., and the relationship information may include: holding a cup in hand, holding a pen in hand, writing on a book with a pen, etc.; behavior information may include drinking, writing, tilting the head, lying down, frowning, etc.; emotional information may include anger, happiness, sadness, etc.

[0069] Analysis based on more comprehensive and detailed multi-dimensional scene understanding information is helpful for subsequent determination of whether there are bad habits, dangerous behaviors, and whether a behavior is maintained for too long, so as to better interact with the user. The details are described in the following embodiments and will not be repeated here.

[0070] In summary, the present application provides a scene understanding method that can continuously acquire multi-dimensional scene information in order to determine the corresponding multi-dimensional scene understanding information. The acquired scene information is more diverse, and the continuous acquisition is conducive to the monitoring of the target monitoring object from a long-term dimension; the analysis results involve four dimensions, making the scene understanding information more comprehensive, which is conducive to better interaction with users in the future, and is conducive to better monitoring the learning situation of the target monitoring object, assisting and assisting its learning, and is conducive to practical application.

[0071] Based on the above embodiment:

[0072] In some embodiments, interacting with a user based on multi-dimensional scene understanding information includes:

[0073] When the behavior recorded in the behavior information continues to appear within a first preset time period, it is determined that the behavior is a learning habit of the target monitored object;

[0074] Determine whether a learning habit is a bad learning habit based on a pre-set habit judgment benchmark;

[0075] If so, the result that the target monitored object has bad learning habits is stored, and / or the prompt module is controlled to prompt the target monitored object that the target monitored object has bad learning habits.

[0076] In this embodiment, considering that students' learning habits are an aspect that parents focus on during the interaction with users, further analysis is performed based on the behavior information included in the multi-dimensional scene understanding information. When the behavior recorded in the behavior information continues to appear within a first preset time period (the first preset time period is flexibly set according to actual needs), it is determined that the behavior is a learning habit of the target monitored object, such as the target monitored object always spinning a pen or singing, sitting tilted on a chair, or keeping eyes close to the book during the learning process; and then, when the learning habit is determined to be a bad learning habit according to the preset habit judgment benchmark, the result of the existence of the bad learning habit is stored so that subsequent users can check the learning situation of the target monitored object; and / or the control prompt module prompts the result of the existence of the bad learning habit, such as prompting that the sitting posture is not correct, the eyes are too close to the book, etc., so that the target monitored object knows and corrects it in time. The prompt module here can be a voice broadcast module in the electronic device to realize the reminder by voice broadcast; it can also be a display module in the electronic device to realize the reminder by display, and there is no special limitation here.

[0077] In addition, the preset habit determination benchmark is essentially a benchmark set by the user in advance for determining which learning habits are bad learning habits. It can specifically include directly set bad learning habits or include the criteria for determining bad learning habits, which are not specifically limited here; and the preset habit determination benchmark can be obtained in advance through the interactive method of voice information or through the interactive method of text information, which are not specifically limited here. It can be understood that the above method makes the solution conducive to flexible adaptation to different learning scenarios of different students, has wide applicability, and is conducive to helping students develop good learning habits and providing assistance for students' learning.

[0078] It should also be noted that the prompt module can also prompt the user of the correct learning habits corresponding to bad learning habits so that the target monitored object can correct and learn.

[0079] In some embodiments, after obtaining the multi-dimensional scene understanding information, the method further includes:

[0080] Store the behavior information and / or emotional information of the target monitoring object at each moment.

[0081] Specifically, by storing the behavior information and / or emotion information at each moment, it is convenient for subsequent users to trace the entire monitoring process.

[0082] In some embodiments, interacting with a user based on multi-dimensional scene understanding information includes:

[0083] According to the preset dangerous behavior judgment criteria, determine whether there is any dangerous behavior in the behavior recorded in the current behavior information;

[0084] If so, the control prompt module prompts the target monitored object to stop performing dangerous behaviors.

[0085] Specifically, the preset dangerous behavior determination criteria are essentially criteria pre-entered and set by the user for determining which behaviors are dangerous behaviors. They may specifically include directly set dangerous behaviors, or may include criteria conditions for determining dangerous behaviors, and are not specifically limited here; and the preset dangerous behavior determination criteria may be obtained in advance in the form of an interactive voice message, or in the form of an interactive text message, and are not specifically limited here.

[0086] Similarly, the prompt module here can be a voice broadcast module or a display module to prompt the target monitored object to stop performing dangerous behaviors.

[0087] In some embodiments, interacting with a user based on multi-dimensional scene understanding information includes:

[0088] According to a preset long-term behavior reminder benchmark, determining whether there is a target behavior in the behavior recorded in the behavior information that has lasted for a second preset time period;

[0089] If so, the control prompt module prompts the target monitored object that the time duration for performing the target behavior has reached a second preset time duration.

[0090] Specifically, the preset long-time behavior reminder benchmark includes at least one target behavior to be reminded and a second preset duration corresponding to the target behavior to be reminded. The second preset duration essentially defines the maximum tolerable time for the target behavior to be maintained. When it is determined that there is a target behavior that lasts for a duration reaching the second preset duration, such as sitting on a chair for 1 hour, the control prompt module prompts the result so that the target monitored object can take the next action accordingly, such as taking a break.

[0091] Furthermore, the preset long-duration behavior reminder benchmark can be obtained in advance through voice information interaction or text information interaction, without any specific limitation here; the specific value of the second preset duration can be flexibly set according to actual needs. Similarly, the prompt module here can be a voice broadcast module or a display module.

[0092] In some embodiments, scene analysis is performed on image information, voice information, and text information according to a preset multimodal analysis strategy to obtain multi-dimensional scene understanding information, including:

[0093] Processing the image information according to N preset image processing algorithms to obtain N basic image information, where N is an integer greater than 1;

[0094] Preprocessing the voice information to obtain corresponding basic voice information;

[0095] Preprocessing the text information to obtain corresponding basic text information;

[0096] Basic image information, basic voice information and basic text information are processed according to the preset multimodal large model to obtain multi-dimensional scene understanding information.

[0097] It should be noted that the preset image processing algorithms include, but are not limited to, various traditional image processing methods, such as edge detection, contour detection, and color conversion; and the N types of basic image information include, but are not limited to, basic image information reflecting contours, basic image information reflecting edges, basic image information reflecting RGB, and basic image information reflecting grayscale. The purpose of setting up N preset image processing algorithms to process image information separately to obtain N types of basic image information is that different preset image processing algorithms target different observation points of image information. The output of N types of basic image information represents consideration of N observation points, which can enrich the observation information of the image and provide richer support for the subsequent processing of the preset multimodal large model.

[0098] In addition, the preprocessing of voice information here may specifically include filtering out noise interference in the collected voice information, etc., which is not specifically limited here; the preprocessing of text information here may specifically include correcting typos in the input text, segmenting and sentence division, etc., which is not specifically limited here.

[0099] It can be seen that the above settings realize the input of multimodal basic image information, basic voice information and basic text information, and combine with the subsequent preset multimodal large model for processing to form a multi-dimensional modal understanding ability to obtain more comprehensive and detailed multi-dimensional scene understanding information.

[0100] In some embodiments, basic image information, basic voice information, and basic text information are processed according to a preset multimodal large model to obtain multi-dimensional scene understanding information, including:

[0101] Input N basic image information as input to a pre-trained image encoding model to extract image features;

[0102] The image features are processed using the image-text alignment module to obtain image features in a text-understandable form;

[0103] Input the basic speech information as input to a pre-trained speech coding model to extract speech features;

[0104] The speech features are processed using the speech-to-text alignment module to obtain speech features in a form that is understandable to text;

[0105] Input the basic text information as input to the pre-trained text encoding model to extract text features;

[0106] Text features, image features in a text-understandable form, and speech features are fed as input to a pre-trained decoding model to obtain multi-dimensional scene understanding information.

[0107] It should be noted that the speech encoding model includes but is not limited to a third network based on the transformer structure; the text encoding model includes but is not limited to a fourth network based on the transformer structure; the image-text alignment module is used to convert image features into features that can be understood by the text, and the speech-text alignment module is used to convert speech features into features that can be understood by the text.

[0108] In some embodiments, N types of basic image information are input as input items to a pre-trained image coding model to extract image features, including:

[0109] Using the first network to extract global features from N kinds of basic image information to obtain global features;

[0110] Using the second network to extract local features from the image sub-regions obtained after each segmentation of the N basic image information to obtain local features;

[0111] The global features and local features are fused using the preset feature fusion strategy to obtain image features.

[0112] It should be noted that the extraction of global features focuses on feature extraction of the entire basic image information, and the extraction of local features focuses on feature extraction of the image sub-regions corresponding to each slice after the basic image information is sliced; the first network here includes but is not limited to a network for global feature extraction based on the transformer structure, and the second network includes but is not limited to a network for local feature extraction based on the transformer structure, and no special limitation is made here; the preset feature fusion strategy here includes but is not limited to the addition of feature dimensions, or the addition or multiplication of values ​​of the same dimension, etc., and no special limitation is made here.

[0113] Exemplarily, in the present application, the feature dimension of the image features extracted by the image coding model can be N×1024×256×256, the feature dimension of the speech features in text-understandable form can be 1×1024, and the feature dimension of the text features in text-understandable form can be 1×1024.

[0114] In some embodiments, text features, image features in a text-understandable form, and speech features are fed as input to a pre-trained decoding model to obtain multi-dimensional scene understanding information, including:

[0115] Inputting text features, image features of a text-understandable form, and speech features as input items into a pre-trained text decoding model to obtain attribute information, relationship information, first behavior information, and first emotion information;

[0116] The text features and the speech features in the text-comprehensible form are input as input items to a pre-trained speech decoding model to obtain the second behavior information and the second emotion information.

[0117] It should be noted that the text decoding model here can be the fifth network set up based on the transformer structure; the speech decoding model here can be the sixth network set up based on the transformer structure, and there is no special limitation here; it can be understood that since the speech information may include sounds such as yawning or laughing of the target monitored object, the output items of the speech decoding model include second behavior information related to the behavior and second emotion information related to the emotion; therefore, the behavior information in the multi-dimensional scene understanding information finally obtained is the union of the first behavior information and the second behavior information, and the emotion information is the union of the first emotion information and the second emotion information.

[0118] In some embodiments, after obtaining the text features and the image features of the text-understandable form, the method further includes:

[0119] The text features and the image features in a text-understandable form are input as input items to a pre-trained image decoding module to obtain an image segmentation map for identifying each object in the image information.

[0120] It should be noted that the image decoding module here can be a seventh network configured based on a cross-attention mechanism, which is not particularly limited here; the image segmentation map here can be an attribute segmentation map and a mask. It can be seen that the above configuration facilitates subsequent interaction with the user based on the image segmentation map.

[0121] In some embodiments, after obtaining the image segmentation map for identifying each object in the image information, the method further includes:

[0122] Annotate each identified object in the image information according to the image segmentation map, and output the annotation results through the display module;

[0123] When a search flag signal indicating that any object given in the annotation result has been parsed is received, a network search is performed in response to the search flag signal to obtain and output a parsing result corresponding to the object.

[0124] Specifically, the display module here can be a display screen; the annotation results can be specifically displayed on the display screen; illustratively, when the user clicks on a labeled object on the display screen, it can be considered that a search mark signal is received, and the corresponding analysis results are output after responding to the search mark signal and searching on the Internet, so that the target monitored object can better understand the current scene, which is especially suitable for lower grade students to perceive the world more realistically, so that they can learn and improve their learning efficiency.

[0125] Please refer to Figure 2 , Figure 2 A schematic diagram of the structure of a scene understanding system provided by the present invention.

[0126] The scene understanding system includes:

[0127] A data acquisition unit 21 is used to continuously acquire multi-dimensional scene information, which includes image information, voice information, and text information;

[0128] A multimodal processing unit 22 is configured to perform scene analysis on the image information, voice information, and text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, wherein the multidimensional scene understanding information includes attribute information representing the settings of each object in the monitoring scene, relationship information representing the association between objects, behavioral information of the target monitoring object, and emotional information of the target monitoring object;

[0129] The comprehensive analysis and interaction unit 23 is used to interact with the user based on multi-dimensional scene understanding information.

[0130] For an introduction to the scene understanding system provided in this application, please refer to the above-mentioned embodiment of the scene understanding method, which will not be repeated here.

[0131] In some embodiments, the comprehensive analysis and interaction unit 23 includes:

[0132] a learning habit determining unit, configured to determine that a behavior recorded in the behavior information is a learning habit of the target monitored subject when the behavior continues to appear within a first preset time period;

[0133] A first judgment unit is configured to judge whether the learning habit is a bad learning habit based on a preset habit judgment benchmark; if so, triggering a first interaction unit;

[0134] The first interaction unit is used to store the result that the target monitoring object has the bad learning habit, and / or control the prompt module to prompt the target monitoring object that the bad learning habit exists.

[0135] In some embodiments, the scene understanding system further includes:

[0136] A storage unit, configured to store the behavior information and / or emotion information of the target monitored object at each moment after obtaining the multi-dimensional scene understanding information;

[0137] In some embodiments, the comprehensive analysis and interaction unit 23 includes:

[0138] The second judgment unit is used to judge whether there is any dangerous behavior in the behavior recorded in the current behavior information according to a preset dangerous behavior judgment benchmark; if so, trigger the second interaction unit;

[0139] The second interaction unit is used to control the prompt module to prompt the target monitored object to stop performing the dangerous behavior.

[0140] In some embodiments, the comprehensive analysis and interaction unit 23 includes:

[0141] a third judgment unit, configured to judge, based on a preset long-time behavior reminder benchmark, whether there is a target behavior in the behavior recorded in the behavior information that has lasted for a second preset time period; if so, trigger the third interaction unit;

[0142] The third interaction unit is used to control the prompt module to prompt the target monitored object that the duration of performing the target behavior has reached the second preset duration.

[0143] In some embodiments, the multimodal processing unit 22 includes:

[0144] an image processing unit, configured to process the image information according to N preset image processing algorithms respectively to obtain N types of basic image information, where N is an integer greater than 1;

[0145] A speech processing unit, configured to pre-process the speech information to obtain corresponding basic speech information;

[0146] A text processing unit, configured to pre-process the text information to obtain corresponding basic text information;

[0147] A multimodal large model processing unit is used to process the basic image information, the basic voice information and the basic text information according to a preset multimodal large model to obtain multi-dimensional scene understanding information.

[0148] In some embodiments, the multimodal large model processing unit includes:

[0149] An image coding unit is used to input N basic image information as input items into a pre-trained image coding model to extract image features;

[0150] a first alignment unit, configured to process the image features using an image-text alignment module to obtain the image features in a text-comprehensible form;

[0151] A speech coding unit, configured to input the basic speech information as an input item into a pre-trained speech coding model to extract speech features;

[0152] a second alignment unit, configured to process the speech features using a speech-to-text alignment module to obtain the speech features in a text-comprehensible form;

[0153] A text encoding unit, configured to input the basic text information as an input item into a pre-trained text encoding model to extract text features;

[0154] A decoding unit is used to input the text features, the image features of the text-understandable form, and the speech features as input items into a pre-trained decoding model to obtain multi-dimensional scene understanding information.

[0155] In some embodiments, the image encoding unit includes:

[0156] A first extraction unit is configured to extract global features from N types of basic image information using a first network to obtain global features;

[0157] A second extraction unit is configured to extract local features from the image sub-regions obtained by slicing the N types of basic image information using a second network to obtain local features;

[0158] The fusion unit is used to fuse the global features and the local features using a preset feature fusion strategy to obtain image features.

[0159] In some embodiments, the decoding unit includes:

[0160] a text decoding unit, configured to input the text features, the image features of the comprehensible form of the text, and the speech features as input items into a pre-trained text decoding model to obtain attribute information, relationship information, first behavior information, and first emotion information;

[0161] a speech decoding unit, configured to input the text features and the speech features of the text-comprehensible form as input items into a pre-trained speech decoding model to obtain second behavior information and second emotion information;

[0162] An image decoding unit is used to input the text features and the image features of the text-understandable form as input items into a pre-trained image decoding module to obtain an image segmentation map for identifying each object in the image information.

[0163] In some embodiments, the scene understanding system further includes:

[0164] a labeling unit, configured to label each identified object in the image information according to the image segmentation map, and output the labeling result through a display module;

[0165] The fourth interaction unit is configured to, upon receiving a search flag signal indicating parsing of any object given in the annotation result, perform an online search in response to the search flag signal to obtain and output a parsing result corresponding to the object.

[0166] Please refer to Figure 3 , Figure 3 This is a structural diagram of an electronic device provided by the present invention.

[0167] The electronic device comprises:

[0168] Memory 31, for storing computer programs;

[0169] The processor 32 is configured to implement the steps of the scene understanding method described above when executing a computer program.

[0170] For an introduction to the electronic device provided in this application, please refer to the above-mentioned embodiment of the scene understanding method, which will not be repeated here.

[0171] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. Relational terms such as first and second are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply that there is any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, the elements limited by the sentence "comprising a" do not exclude the presence of other identical elements in the process, method, article or equipment including the elements.

[0172] The above description of the disclosed embodiments will enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is to be construed in the widest manner consistent with the principles and novel features disclosed herein.

Claims

1. A scene understanding method, characterized in that: include: Continuously acquiring multi-dimensional scene information, wherein the multi-dimensional scene information includes image information, voice information, and text information; Performing scene analysis on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, wherein the multidimensional scene understanding information includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between the objects, behavior information of the target monitoring object, and emotional information of the target monitoring object; Interact with the user based on the multi-dimensional scene understanding information.

2. The scene understanding method according to claim 1, wherein: Interacting with the user based on the multi-dimensional scene understanding information includes: When the behavior recorded in the behavior information continues to appear within a first preset time period, determining that the behavior is a learning habit of the target monitored object; Determining whether the learning habit is a bad learning habit based on a preset habit determination benchmark; If so, the result that the target monitoring object has the bad learning habit is stored, and / or the prompt module is controlled to prompt the target monitoring object that the bad learning habit exists.

3. The scene understanding method according to claim 1, wherein: After obtaining multi-dimensional scene understanding information, it also includes: Store the behavior information and / or emotion information of the target monitoring object at each moment.

4. The scene understanding method according to claim 1, wherein: Interacting with the user based on the multi-dimensional scene understanding information includes: According to the preset dangerous behavior judgment criteria, determine whether there is any dangerous behavior in the behavior recorded in the current behavior information; If so, the control prompt module prompts the target monitored object to stop performing the dangerous behavior.

5. The scene understanding method according to claim 1, wherein: Interacting with the user based on the multi-dimensional scene understanding information includes: According to a preset long-time behavior reminder benchmark, determining whether there is a target behavior in the behavior recorded in the behavior information that has lasted for a second preset time period; If so, the control prompt module prompts the target monitoring object that the time duration for performing the target behavior has reached the second preset time duration.

6. The scene understanding method according to any one of claims 1 to 5, wherein: Performing scene analysis on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multi-dimensional scene understanding information, including: Processing the image information according to N preset image processing algorithms respectively to obtain N basic image information, where N is an integer greater than 1; Preprocessing the voice information to obtain corresponding basic voice information; Preprocessing the text information to obtain corresponding basic text information; The basic image information, the basic voice information and the basic text information are processed according to a preset multimodal large model to obtain multi-dimensional scene understanding information.

7. The scene understanding method according to claim 6, wherein: The basic image information, the basic voice information, and the basic text information are processed according to a preset multimodal large model to obtain multi-dimensional scene understanding information, including: Input N basic image information as input to a pre-trained image encoding model to extract image features; Processing the image features using an image-text alignment module to obtain the image features in a text-comprehensible form; Inputting the basic speech information as an input item into a pre-trained speech coding model to extract speech features; Processing the speech features using a speech-to-text alignment module to obtain the speech features in a text-comprehensible form; Inputting the basic text information as input into a pre-trained text encoding model to extract text features; The text features, the image features of the text-understandable form, and the speech features are input as input items to a pre-trained decoding model to obtain multi-dimensional scene understanding information.

8. The scene understanding method according to claim 7, wherein: Input N basic image information as input to a pre-trained image encoding model to extract image features, including: Using the first network to extract global features from N kinds of basic image information to obtain global features; Using a second network, local feature extraction is performed on the image sub-regions obtained after each of the N types of basic image information is fragmented to obtain local features; The global features and the local features are fused using a preset feature fusion strategy to obtain image features.

9. The scene understanding method according to claim 7, wherein: The text features, the image features of the text-understandable form, and the speech features are input into a pre-trained decoding model to obtain multi-dimensional scene understanding information, including: Inputting the text features, the image features of the understandable form of the text, and the speech features as input items into a pre-trained text decoding model to obtain attribute information, relationship information, first behavior information, and first emotion information; The text features and the speech features of the text-understandable form are input as input items to a pre-trained speech decoding model to obtain second behavior information and second emotion information.

10. The scene understanding method according to claim 7, wherein: After obtaining the text features and the image features of the understandable form of the text, the method further includes: The text features and the image features in the understandable form of the text are input as input items to a pre-trained image decoding module to obtain an image segmentation map for identifying each object in the image information.

11. The scene understanding method according to claim 10, wherein: After obtaining the image segmentation map for identifying each object in the image information, the method further includes: Annotating each identified object in the image information according to the image segmentation map, and outputting the annotation result through a display module; When a search flag signal indicating that any object given in the annotation result has been parsed is received, a network search is performed in response to the search flag signal to obtain and output a parsing result corresponding to the object.

12. A scene understanding system, characterized in that: include: A data acquisition unit, configured to continuously acquire multi-dimensional scene information, wherein the multi-dimensional scene information includes image information, voice information, and text information; a multimodal processing unit, configured to perform scene analysis on the image information, the voice information, and the text information according to a preset multimodal analysis strategy to obtain multidimensional scene understanding information, wherein the multidimensional scene understanding information includes attribute information representing the setting of each object in the monitoring scene, relationship information representing the association between the objects, behavioral information of the target monitoring object, and emotional information of the target monitoring object; A comprehensive analysis and interaction unit is used to interact with the user based on the multi-dimensional scene understanding information.

13. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the scene understanding method according to any one of claims 1 to 11 when executing the computer program.