Methods and devices for processing sensed information, and household appliances
Patent Information
- Application Number
- CN202210447091.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-26
- Publication Date
- 2026-03-06
- Estimated Expiration
- 2042-04-26
AI Technical Summary
[0003]本申请实施例提供了一种感知信息的处理方法及装置、家电设备,以至少解决人工智能助理在感知用户场景时,有时通过单一模态的感知传感器无法准确识别场景,但是如果同时开启多种模态的感知传感器又会造成资源浪费的技术问题
[0016]在本申请实施例中,采用获取第一感知传感器感知到的第一感知信息,其中,第一感知传感器为单一模态的感知传感器,第一感知信息用于识别目标场景;确定利用第一感知信息识别目标场景对应的置信度;在置信度满足预设条件的情况下,控制第二感知传感器开启,并获取第二感知传感器感知到的第二感知信息;利用第一感知信息和第二感知信息识别目标场景的方式,通过在单一模态的感知传感器开启后,通过感知结果判断,若感知信息缺失,则增加其他模态的感知传感器,以补齐单一模态传感器感知信息的缺失,达到了解决单一模态的感知传感器无法准确识别场景的问题,同时,可以按需开启某一模态的感知传感器,避免了所有场景下都同时需要开启多模态感知传感器的目的,从而实现了提升了AI助理的多模态感知能力,在不降低用户体验的前提下节约了机器运行资源的技术效果,进而解决了人工智能助理在感知用户场景时,有时通过单一模态的感知传感器无法准确识别场景,但是如果同时开启多种模态的感知传感器又会造成资源浪费技术问题。
Smart Images

Figure CN114913609B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of scene recognition technology, and more specifically, to a method and apparatus for processing perceived information, a home appliance, and a computer-readable storage medium. Background Technology
[0002] When artificial intelligence (AI) assistants perceive user scenarios, sometimes relying on a single modality of perception (such as microphone acoustics) is insufficient to accurately identify the scenario; for example, expecting to identify whether a user has a cold or fever through sound. However, activating multiple perception capabilities simultaneously (such as microphone acoustics and vision) is a waste of machine resources. Summary of the Invention
[0003] This application provides a method and apparatus for processing sensing information, as well as a home appliance, to at least solve the technical problem that when an artificial intelligence assistant perceives a user scenario, sometimes it cannot accurately identify the scenario using a single-modality sensing sensor, but if multiple-modality sensing sensors are activated simultaneously, it will cause a waste of resources.
[0004] According to one aspect of the embodiments of this application, a method for processing sensing information is provided, comprising: acquiring first sensing information sensed by a first sensing sensor, wherein the first sensing sensor is a single-mode sensing sensor, and the first sensing information is used to identify a target scene; determining the confidence level corresponding to identifying the target scene using the first sensing information; when the confidence level meets a preset condition, controlling a second sensing sensor to turn on and acquiring second sensing information sensed by the second sensing sensor; and identifying the target scene using the first sensing information and the second sensing information.
[0005] Optionally, before controlling the second sensing sensor to turn on, the above method further includes: determining whether the confidence level is lower than a first preset threshold; if the confidence level is lower than the first preset threshold, determining that the confidence level meets the preset condition; if the confidence level is higher than the first preset threshold, determining that the confidence level does not meet the preset condition.
[0006] Optionally, the first perception information includes audio information, and the second perception sensor is a camera. Acquiring the second perception information perceived by the second perception sensor includes: receiving multiple image frames acquired by the camera; performing image detection on the multiple image frames to determine the action features corresponding to each image frame; and determining the target image frame containing the target action features as the second perception information, wherein the target action features are correlated with the audio information.
[0007] Optionally, the target action feature includes: a first action feature and a second action feature, wherein the first action feature is an action feature corresponding to the audio information; and the second action feature is an action feature that is related to the first action feature and the audio information.
[0008] Optionally, identifying a target scene using first and second perceptual information includes: determining a first action keyword corresponding to the audio information and a second action keyword corresponding to the target action feature; determining the matching degree between the first action keyword and the second action keyword; and if the matching degree is higher than a second preset threshold, using the target scene identified based on the audio information as the final scene identification result.
[0009] Optionally, the above method further includes: obtaining the recognition result of the target home appliance on the target scene, wherein the target home appliance is a home appliance located within the same preset range as the home appliance where the first sensing sensor and the second sensing sensor are located; if the recognition result is the same as the target scene, it is determined that the target scene has been successfully recognized.
[0010] Optionally, the second sensing sensor includes: a single-modality sensing sensor or a multi-modality sensing sensor; when the second sensing sensor includes a multi-modality sensing sensor, activating the second sensing sensor includes: determining the type of the target scene to be identified; searching the database for the type of the second sensing sensor corresponding to the type of the target scene; and activating the second sensing sensor of the type corresponding to the type of the target scene.
[0011] Optionally, after identifying the target scene using the first and second perception information, the method further includes: generating prompt information, wherein the prompt information is used to prompt the target object to perform an operation corresponding to the target scene; and controlling the home appliance to perform an operation corresponding to the target scene.
[0012] According to another aspect of the embodiments of this application, a processing device for sensing information is also provided, comprising: a first acquisition module, configured to acquire first sensing information perceived by a first sensing sensor, wherein the first sensing sensor is a single-mode sensing sensor, and the first sensing information is used to identify a target scene; a first determination module, configured to determine a first confidence level corresponding to identifying the target scene using the first sensing information; a processing module, configured to control a second sensing sensor to turn on and acquire second sensing information perceived by the second sensing sensor when the first confidence level meets a preset condition; and an identification module, configured to identify the target scene using the first sensing information and the second sensing information.
[0013] According to another aspect of the embodiments of this application, a home appliance is also provided, including: a controller and a plurality of sensing sensors, wherein the plurality of sensing sensors are used to sense sensing information, wherein the sensing information is used to identify a target scene; the controller is connected to the plurality of sensing sensors and is used to execute the above-described sensing information processing method.
[0014] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which includes a stored program, wherein the program executes the above-described method for processing perceived information when it is run.
[0015] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described method for processing perceived information through the computer program.
[0016] In this embodiment, the method involves acquiring first sensing information from a first sensing sensor, wherein the first sensing sensor is a single-modality sensing sensor, and the first sensing information is used to identify a target scene; determining the confidence level corresponding to identifying the target scene using the first sensing information; and controlling the second sensing sensor to turn on when the confidence level meets a preset condition, and acquiring the second sensing information from the second sensing sensor. This method of identifying the target scene using the first and second sensing information solves the problem of single-modality sensing sensors being unable to accurately identify scenes. Furthermore, it allows for the activation of a single-modality sensing sensor as needed, avoiding the need to activate multiple modal sensing sensors simultaneously in all scenarios. This enhances the multimodal perception capability of the AI assistant and saves machine resources without compromising user experience. Ultimately, it addresses the technical problem that sometimes a single-modality sensing sensor cannot accurately identify a user's scene, while activating multiple modal sensing sensors simultaneously would waste resources. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the hardware environment for a method of processing perceived information according to an embodiment of this application;
[0020] Figure 2 This is a flowchart of a method for processing perceived information according to an embodiment of this application;
[0021] Figure 3 This is a structural block diagram of a sensory information processing device according to an embodiment of this application;
[0022] Figure 4 This is a structural block diagram of a household appliance according to an embodiment of this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0025] According to one aspect of the embodiments of this application, a method for processing sensing information is provided. This method for processing sensing information is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned method for processing sensing information can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0026] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0027] Figure 2 This is a flowchart of a method for processing perceived information according to an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps:
[0028] Step S202: Obtain first perception information perceived by the first perception sensor, wherein the first perception sensor is a single-mode perception sensor, and the first perception information is used to identify the target scene.
[0029] According to an optional embodiment of this application, the first sensing sensor is a single-modal sensing sensor, such as a microphone for acoustic recognition.
[0030] When the first sensing sensor is a microphone, the first sensing information is the collected sound information.
[0031] Step S204: Determine the confidence level corresponding to the target scene identified using the first perception information.
[0032] First, let's explain the three related concepts involved:
[0033] Point estimation uses sample statistics to estimate population parameters. Since the sample statistics are values at a point on the number line, the estimation results are also represented by a single point value.
[0034] Interval estimation, based on point estimation, provides a range of estimates for population parameters. This range is usually obtained by adding or subtracting the estimation difference from the sample statistic, and the resulting range is the confidence interval.
[0035] Confidence level, also known as reliability, confidence level, or confidence coefficient, refers to the accuracy (reliability) of an interval estimate. When estimating population parameters through sampling, the conclusion is always uncertain due to the randomness of the sample. Therefore, a probabilistic statement, namely interval estimation in mathematical statistics, is used. The confidence level measures the probability that the estimated value falls within a certain allowable error range from the population parameter.
[0036] Confidence level (confidence coefficient) is based on the frequency interpretation of probability, assuming that the estimated θ is an unknown but definite quantity, while the sample x is random. Whether the interval [A(x), B(x)] contains the estimated θ depends on the sample x. Therefore, the interval [A(x), B(x)] can only contain the unknown θ with a certain probability. For different θ, the value of π(θ) can be different, and the minimum value of π(θ) for different θ, 1-α (0 < α < 1), is called the confidence coefficient of the interval [A(x), B(x)]. Correspondingly, the interval [A(x), B(x)] is called a confidence interval of θ.
[0037] Therefore, the confidence coefficient can be interpreted as the degree to which one can be confident about the inference that "the interval [A(x), B(x)] contains θ", and the degree of confidence is represented by the confidence coefficient.
[0038] Specifically, in the embodiments of this application, the confidence level in step S204 refers to the degree of credibility of the probability of identifying the target scene using the first perception information within a confidence interval. In this step, firstly, the confidence interval for identifying the target scene using the first perception information is determined, and then the degree of credibility of the probability of identifying the target scene using the first perception information within this confidence interval is calculated, thus obtaining the confidence level of identifying the target scene using the first perception information.
[0039] The target scenario mentioned above can be a sound or an action, such as the sound and action of a user "sneezing".
[0040] Step S206: If the confidence level meets the preset conditions, control the second sensing sensor to turn on and acquire the second sensing information sensed by the second sensing sensor.
[0041] It should be noted that the second sensing sensor is a different type of sensor than the first sensor.
[0042] Step S208: Identify the target scene using the first and second perception information.
[0043] Through the above steps, after a single-modal perception sensor is activated, the perception results are used to determine if there is missing perception information. If so, other modal perception sensors are added to compensate for the lack of perception information from the single-modal sensor. This solves the problem that a single-modal perception sensor cannot accurately identify the scene. At the same time, a certain modal perception sensor can be activated as needed, avoiding the need to activate multiple modal perception sensors in all scenarios. This improves the multimodal perception capability of the AI assistant and saves machine operating resources without reducing the user experience.
[0044] According to an optional embodiment of this application, before executing step S206 to activate the second sensing sensor, it is also necessary to determine whether the confidence level is lower than a first preset threshold; if the confidence level is lower than the first preset threshold, it is determined that the confidence level meets the preset condition; if the confidence level is higher than the first preset threshold, it is determined that the confidence level does not meet the preset condition.
[0045] In an optional embodiment, when identifying whether a user has sneezed, a microphone is used to acoustically identify whether a sneezing sound exists. After the audio information is collected by the microphone, an AI assistant is used to determine whether the confidence level of the collected audio information is sufficient to confirm the existence of a sneezing sound. Specifically, if the confidence level of the collected audio information is lower than a preset threshold, it is determined that the collected audio information cannot be used to confirm the existence of a sneezing sound (at this time, the confidence level meets the preset condition). If the confidence level of the collected audio information is higher than the preset threshold, it is determined that the collected audio information can be used to confirm the existence of a sneezing sound (at this time, the confidence level does not meet the preset condition).
[0046] According to another optional embodiment of this application, the first sensing information includes: audio information, and the second sensing sensor is a camera. Step S206 is performed to obtain the second sensing information perceived by the second sensing sensor, including the following steps: receiving multiple image frames acquired by the camera; performing image detection on the multiple image frames to determine the action features corresponding to each image frame; determining the target image frame containing the target action features in the multiple image frames as the second sensing information, wherein the target action features are related to the audio information.
[0047] Taking the scenario of identifying whether a user is "sneezing" mentioned above as an example, the first sensing information is audio information. When the audio information cannot be used to identify whether a user is "sneezing", the second sensing sensor (i.e., the camera) is activated. Correspondingly, the second sensing information is the image frames captured by the camera. It should be noted that the image frames captured by the second sensing sensor are related to the audio information captured by the first sensing sensor, and both can extract some features that characterize the action of "sneezing".
[0048] When performing step S206 to obtain the second perception information perceived by the second perception sensor, multiple image frames captured by the camera are detected, and image frames containing specific target action features (e.g., specific action features corresponding to "sneezing") are used as the second perception information corresponding to the second perception sensor.
[0049] As an optional embodiment, the target action feature includes: a first action feature and a second action feature, wherein the first action feature is an action feature corresponding to the audio information; and the second action feature is an action feature that is associated with the first action feature and the audio information.
[0050] Taking the identification of the target scene corresponding to "sneezing" as an example, the camera is used to identify whether there are specific action features corresponding to "sneezing" (i.e. the first action feature mentioned above) or actions after sneezing (such as wiping the nose, i.e. the second action feature mentioned above).
[0051] In some optional embodiments of this application, the execution step S208, which uses the first perception information and the second perception information to identify the target scene, can be achieved by the following method: determining the first action keyword corresponding to the audio information and the second action keyword corresponding to the target action feature; determining the matching degree between the first action keyword and the second action keyword; and if the matching degree is higher than a second preset threshold, taking the target scene identified based on the audio information as the final scene identification result.
[0052] In this step, the image frames collected by the second sensing sensor are used to supplement the audio information collected by the first sensing sensor. Then, the AI assistant makes a comprehensive decision to confirm whether the user has sneezed.
[0053] By using the above method, a specific mode of sensing sensor can be activated as needed, thus avoiding the waste of resources caused by activating multiple modes of sensing sensors simultaneously in all scenarios.
[0054] In some optional embodiments of this application, the recognition result of the target home appliance on the target scene can also be obtained, wherein the target home appliance is a home appliance located within the same preset range as the home appliance where the first sensing sensor and the second sensing sensor are located; if the recognition result is the same as the target scene, it is determined that the target scene has been successfully recognized.
[0055] To improve the accuracy of target scene recognition, the recognition results of other home appliances located in the same preset area can also be obtained. If the recognition results of other home appliances on the target scene are the same as the recognition results of the target scene using the first sensing sensor combined with the second sensing sensor, the target scene is successfully recognized.
[0056] For example, the system can identify whether a user is sneezing by using a standing air conditioner in the living room. After the acoustic and video sensors on the standing air conditioner identify the user's sneezing action, it can then obtain the recognition result of the user's sneezing action from the air conditioner in the bedroom. If the air conditioner in the bedroom also identifies the user's sneezing action, it can be determined that the recognition result of the standing air conditioner is accurate.
[0057] By using the above method and verifying the recognition results of the first and second sensing sensors, the technical effect of improving the accuracy of home appliances in recognizing target scenes can be achieved.
[0058] In some optional embodiments of this application, the second sensing sensor includes: a single-modality sensing sensor or a multi-modality sensing sensor; when the second sensing sensor includes a multi-modality sensing sensor, activating the second sensing sensor includes: determining the type of the target scene to be identified; searching the database for the type of the second sensor corresponding to the type of the target scene; and activating the second sensing sensor of the type corresponding to the type of the target scene.
[0059] In the embodiment mentioned above for identifying whether a user is "sneezing," the second sensing sensor is a camera used to acquire image frames, and is a single-modality sensing sensor. In applications where multiple sensing sensors are needed to identify various types of sensory information, the second sensing sensor may include multi-modality sensing sensors.
[0060] Because different scenarios require different types of sensors—for example, recognizing a "sneezing" scenario is better done with an image sensor capturing the user's image, while recognizing a "sweating" scenario is better done with a temperature sensor capturing the user's body temperature—it's possible to pre-store the mapping between scenario types and sensor types in a database. When a sensor is activated, the database is used to find the corresponding sensor type based on the target scenario. Selecting the appropriate sensor for each scenario not only improves the accuracy of scenario recognition but also avoids the waste of resources caused by blindly activating multiple sensors.
[0061] In some other optional embodiments of this application, after step S208 identifies the target scene using the first perception information and the second perception information, a prompt message is generated, wherein the prompt message is used to prompt the target object to perform an operation corresponding to the target scene.
[0062] Taking the scenario of recognizing whether a user is "sneezing" mentioned above as an example, after confirming that the user has "sneezed," the system can remind the user through voice interaction to close the window or increase the indoor air conditioning temperature.
[0063] As an optional embodiment, after performing step S208 to identify the target scene using the first and second perception information, the home appliance can also be controlled to perform an operation corresponding to the target scene.
[0064] In the example above, after confirming that the user has sneezed, the system uses voice interaction to remind the user to close the window or raise the indoor air conditioning temperature. It can also directly control the air conditioning to raise the temperature.
[0065] The method provided in this application improves the multimodal perception capability of AI assistants, solves the problem of not being able to accurately identify a scene with a single modal sensor, and enables the activation of a certain modal perception sensor as needed, avoiding the need to activate multimodal perception sensors in all scenarios at the same time, thus achieving the technical effect of saving machine operating resources without reducing user experience.
[0066] Figure 3 This is a structural block diagram of a sensing information processing device according to an embodiment of this application, such as... Figure 3 As shown, the device includes:
[0067] The first acquisition module 30 is used to acquire first perception information perceived by the first perception sensor, wherein the first perception sensor is a single-mode perception sensor, and the first perception information is used to identify the target scene.
[0068] According to an optional embodiment of this application, the first sensing sensor is a single-modal sensing sensor, such as a microphone for acoustic recognition. When the first sensing sensor is a microphone, the first sensing information is the acquired sound information.
[0069] The first determining module 32 is used to determine the confidence level corresponding to the target scene identified using the first perception information.
[0070] The target scenario mentioned above can be a sound or an action, such as the sound and action of a user "sneezing".
[0071] The processing module 34 is used to control the second sensing sensor to turn on and acquire the second sensing information sensed by the second sensing sensor when the confidence level meets the preset conditions.
[0072] It should be noted that the second sensing sensor is a different type of sensor than the first sensor.
[0073] The recognition module 36 is used to identify the target scene using the first perception information and the second perception information.
[0074] It should be noted that, Figure 3 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.
[0075] According to an optional embodiment of this application, the above-described apparatus further includes:
[0076] The judgment module is used to determine whether the confidence level is lower than a first preset threshold before the second sensing sensor is activated;
[0077] The second determining module is used to determine that the confidence level meets the preset conditions when the confidence level is lower than the first preset threshold.
[0078] The third determining module is used to determine that the confidence level does not meet the preset conditions when the confidence level is higher than the first preset threshold.
[0079] In an optional embodiment, when identifying whether a user has sneezed, a microphone is used to acoustically identify whether a sneezing sound exists. After the audio information is collected by the microphone, an AI assistant is used to determine whether the confidence level of the collected audio information is sufficient to confirm the existence of a sneezing sound. Specifically, if the confidence level of the collected audio information is lower than a preset threshold, it is determined that the collected audio information cannot be used to confirm the existence of a sneezing sound (at this time, the confidence level meets the preset condition). If the confidence level of the collected audio information is higher than the preset threshold, it is determined that the collected audio information can be used to confirm the existence of a sneezing sound (at this time, the confidence level does not meet the preset condition).
[0080] According to another optional embodiment of this application, the first perception information includes: audio information, the second perception sensor is a camera, and the above-mentioned processing module 34 is further configured to receive multiple image frames acquired by the camera; perform image detection on the multiple image frames to determine the action features corresponding to each image frame; and determine the target image frame containing the target action features in the multiple image frames as the second perception information, wherein the target action features are related to the audio information.
[0081] Taking the scenario of identifying whether a user is "sneezing" mentioned above as an example, the first sensing information is audio information. When the audio information cannot be used to identify whether a user is "sneezing", the second sensing sensor (i.e., the camera) is activated. Correspondingly, the second sensing information is the image frames captured by the camera. It should be noted that the image frames captured by the second sensing sensor are related to the audio information captured by the first sensing sensor, and both can extract some features that characterize the action of "sneezing".
[0082] As an optional embodiment, the target action feature includes: a first action feature and a second action feature, wherein the first action feature is an action feature corresponding to the audio information; and the second action feature is an action feature that is associated with the first action feature and the audio information.
[0083] Taking the identification of the target scene corresponding to "sneezing" as an example, the camera is used to identify whether there are specific action features corresponding to "sneezing" (i.e. the first action feature mentioned above) or actions after sneezing (such as wiping the nose, i.e. the second action feature mentioned above).
[0084] In some optional embodiments of this application, the identification module 36 includes:
[0085] The first determining unit is configured to determine the first action keyword corresponding to the audio information and the second action keyword corresponding to the target action feature; the second determining unit is configured to determine the matching degree between the first action keyword and the second action keyword; the processing unit is configured to take the target scene obtained based on the audio recognition as the final scene recognition result when the matching degree is higher than a second preset threshold.
[0086] The AI assistant uses image frames collected by the second sensing sensor to supplement the audio information collected by the first sensing sensor, and then makes a comprehensive decision to confirm whether the user has sneezed.
[0087] By activating a specific mode of sensing sensor on demand, the problem of wasting resources caused by activating multiple modes of sensing sensors simultaneously in all scenarios is avoided.
[0088] According to an optional embodiment of this application, the above-mentioned device further includes: a second acquisition module, used to acquire the recognition result of the target home appliance on the target scene, wherein the target home appliance is a home appliance located within the same preset range as the home appliance where the first sensing sensor and the second sensing sensor are located; and a fourth determination module, used to determine that the target scene has been successfully recognized when the recognition result is the same as the target scene.
[0089] To improve the accuracy of target scene recognition, the recognition results of other home appliances located in the same preset area can also be obtained. If the recognition results of other home appliances on the target scene are the same as the recognition results of the target scene using the first sensing sensor combined with the second sensing sensor, the target scene is successfully recognized.
[0090] For example, the system can identify whether a user is sneezing by using a standing air conditioner in the living room. After the acoustic and video sensors on the standing air conditioner identify the user's sneezing action, it can then obtain the recognition result of the user's sneezing action from the air conditioner in the bedroom. If the air conditioner in the bedroom also identifies the user's sneezing action, it can be determined that the recognition result of the standing air conditioner is accurate.
[0091] By verifying the recognition results of the first and second sensing sensors, the technical effect of improving the accuracy of home appliances in recognizing target scenes can be achieved.
[0092] According to another optional embodiment of this application, the second sensing sensor includes: a single-modality sensing sensor or a multi-modality sensing sensor; when the second sensing sensor includes a multi-modality sensing sensor, the processing module 34 is further configured to determine the type of the target scene to be identified; search the database for the type of the second sensing sensor corresponding to the type of the target scene; and activate the second sensing sensor of the type corresponding to the type of the target scene.
[0093] In the embodiment mentioned above for identifying whether a user is "sneezing," the second sensing sensor is a camera used to acquire image frames, and is a single-modality sensing sensor. In applications where multiple sensing sensors are needed to identify various types of sensory information, the second sensing sensor may include multi-modality sensing sensors.
[0094] Because different scenarios require different types of sensors—for example, recognizing a "sneezing" scenario is better done with an image sensor capturing the user's image, while recognizing a "sweating" scenario is better done with a temperature sensor capturing the user's body temperature—it's possible to pre-store the mapping between scenario types and sensor types in a database. When a sensor is activated, the database is used to find the corresponding sensor type based on the target scenario. Selecting the appropriate sensor for each scenario not only improves the accuracy of scenario recognition but also avoids the waste of resources caused by blindly activating multiple sensors.
[0095] In some other optional embodiments of this application, the above-mentioned apparatus further includes: a generation module, used to generate prompt information, wherein the prompt information is used to prompt the target object to perform an operation corresponding to the target scene.
[0096] Taking the scenario of recognizing whether a user is "sneezing" mentioned above as an example, after confirming that the user has "sneezed," the system can remind the user through voice interaction to close the window or increase the indoor air conditioning temperature.
[0097] As an optional embodiment, the above-mentioned device further includes: a control module for controlling the home appliance to perform an operation corresponding to the target scenario.
[0098] In the example above, after confirming that the user has sneezed, the system uses voice interaction to remind the user to close the window or raise the indoor air conditioning temperature. It can also directly control the air conditioning to raise the temperature.
[0099] The device provided in this application enhances the multimodal perception capability of AI assistants, solves the problem of not being able to accurately identify a scene with a single-modal sensor, and enables the activation of a certain mode of perception sensor as needed, avoiding the need to activate multimodal perception sensors in all scenarios at the same time, thus achieving the technical effect of saving machine operating resources without reducing user experience.
[0100] Figure 4 This is a structural block diagram of a household appliance according to an embodiment of this application, such as... Figure 4 The device includes a controller 40 and multiple sensing sensors 42, wherein the multiple sensing sensors 42 are used to sense sensing information, and the sensing information is used to identify the target scene; the controller 40 is connected to the multiple sensing sensors 42 and is used to execute the above-mentioned sensing information processing method.
[0101] In the embodiments provided in this application, the home appliances include, but are not limited to, air conditioners.
[0102] Instructions are required. Figure 4Preferred embodiments of the shown examples can be found in [reference needed]. Figure 2 The relevant descriptions of the embodiments shown will not be repeated here.
[0103] This application also provides a computer-readable storage medium, which includes a stored program, wherein the program executes the above-described method for processing perceived information when it runs.
[0104] The aforementioned computer-readable storage medium is used to store a program that performs the following functions: acquiring first sensing information perceived by a first sensing sensor, wherein the first sensing sensor is a single-mode sensing sensor, and the first sensing information is used to identify a target scene; determining the confidence level corresponding to identifying the target scene using the first sensing information; controlling the second sensing sensor to turn on when the confidence level meets a preset condition, and acquiring second sensing information perceived by the second sensing sensor; and identifying the target scene using the first sensing information and the second sensing information.
[0105] This application also provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to execute the above-mentioned method for processing perceived information through the computer program.
[0106] The processor is configured to perform the following functions via a computer program: acquire first sensing information perceived by a first sensing sensor, wherein the first sensing sensor is a single-mode sensing sensor and the first sensing information is used to identify a target scene; determine the confidence level corresponding to the identification of the target scene using the first sensing information; when the confidence level meets a preset condition, control the second sensing sensor to turn on and acquire the second sensing information perceived by the second sensing sensor; and identify the target scene using the first sensing information and the second sensing information.
[0107] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A processing method of perception information, characterized by, The method comprises: obtaining first perception information perceived by a first perception sensor, wherein the first perception sensor is a single-modal perception sensor, and the first perception information is used to identify a target scene; determining a confidence degree of identifying the target scene corresponding to the first perception information; in a case where the confidence degree meets a preset condition, controlling a second perception sensor to be turned on, and obtaining second perception information perceived by the second perception sensor; identifying the target scene by using the first perception information and the second perception information; The method further comprises: obtaining an identification result of the target scene by other home appliances located in the same preset area; if the identification result is the same as the target scene, determining that the target scene is successfully identified; wherein, in a case where the first perception information comprises audio information, and the second perception information comprises a target image frame containing a target action feature, identifying the target scene by using the first perception information and the second perception information comprises: determining a first action keyword corresponding to the audio information and a second action keyword corresponding to the target action feature; determining a matching degree of the first action keyword and the second action keyword; in a case where the matching degree is higher than a second preset threshold, taking the target scene identified according to the audio information as a final scene identification result, wherein the target action feature has a correlation with the audio information.
2. The method of claim 1, wherein, Before controlling the second perception sensor to be turned on, the method further comprises: determining whether the confidence degree is lower than a first preset threshold; if the confidence degree is lower than the first preset threshold, determining that the confidence degree meets the preset condition; if the confidence degree is higher than the first preset threshold, determining that the confidence degree does not meet the preset condition.
3. The method of claim 1 or 2, wherein: the first perception information comprises audio information, the second perception sensor is a camera, and the obtaining of the second perception information perceived by the second perception sensor comprises: receiving a plurality of image frames obtained by the camera; performing image detection on the plurality of image frames to determine an action feature corresponding to each image frame; determining a target image frame containing a target action feature from the plurality of image frames as the second perception information.
4. The method of claim 3, wherein, The target action feature comprises a first action feature and a second action feature, wherein: the first action feature is an action feature corresponding to the audio information; and the second action feature is an action feature having a correlation with the first action feature and the audio information.
5. The method of claim 1, wherein: the second perception sensor comprises a single-modal perception sensor or a plurality of modal perception sensors; in a case where the second perception sensor comprises a plurality of modal perception sensors, the controlling of the second perception sensor to be turned on comprises: determining a type of the target scene to be identified; finding a type of the second perception sensor corresponding to the type of the target scene from a database; and turning on the second perception sensor of the type corresponding to the type of the target scene.
6. The method of claim 1, wherein, After the target scene is identified by using the first perception information and the second perception information, the method further includes: generating prompt information, wherein the prompt information is used to prompt a target object to perform an operation corresponding to the target scene; controlling an electrical appliance to perform an operation corresponding to the target scene.
7. A processing device of perception information, characterized by, comprise: a first acquisition module, configured to acquire first perception information perceived by a first perception sensor, wherein the first perception sensor is a single-modal perception sensor, and the first perception information is used to identify a target scene; a first determination module, configured to determine a confidence degree of identifying the target scene by using the first perception information; a processing module, configured to, in a case where the confidence degree meets a preset condition, control a second perception sensor to be turned on, and acquire second perception information perceived by the second perception sensor; an identification module, configured to identify the target scene by using the first perception information and the second perception information; the apparatus further comprises a second acquisition module, configured to acquire an identification result of the target scene by other electrical appliances located in a same preset area; and if the identification result is the same as the target scene, it is determined that the target scene is successfully identified. In a case where the first perception information comprises audio information, and the second perception information comprises a target image frame containing a target action feature, the identification module is further configured to determine a first action keyword corresponding to the audio information and a second action keyword corresponding to the target action feature; determine a matching degree of the first action keyword and the second action keyword; and in a case where the matching degree is higher than a second preset threshold, take the target scene identified according to the audio information as a final scene identification result, wherein the target action feature has a correlation with the audio information.
8. An electric home appliance characterized by comprising: comprise: a controller and a plurality of perception sensors, wherein the plurality of perception sensors are configured to perceive perception information, wherein the perception information is used to identify a target scene; the controller, connected with the plurality of perception sensors, is configured to perform the processing method of the perception information according to any one of claims 1 to 6.
9. A computer readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program performs the processing method of the perception information according to any one of claims 1 to 6 when running. 10.An electronic device comprising a memory and a processor, the electronic device characterized by, The memory stores a computer program, and the processor is configured to perform the processing method of the perception information according to any one of claims 1 to 6 by using the computer program.
Citation Information
Patent Citations
Context recognition method and device, terminal and computer readable storage medium
CN110516760A
Temperature adjusting method and device based on image recognition, electronic equipment and medium
CN113705467A
Control method and device of intelligent equipment
CN113791557A