Tipping methods, devices, electronic devices and readable storage media

By acquiring images from different angles through a dual-camera system and using stereo parallax technology to calculate the probability of learning states, the problem of low accuracy in detecting children's learning in existing technologies has been solved, achieving more efficient learning supervision.

CN115050054BActive Publication Date: 2026-03-06VIVO MOBILE COMM CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202210763061.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-30
Publication Date
2026-03-06
Estimated Expiration
2042-06-30

AI Technical Summary

Technical Problem

Current methods for detecting whether a child is studying have low accuracy.

Method used

A dual-camera system is used to acquire images from different angles, extract stereoscopic feature information using stereo parallax technology, calculate the probability value of the target object being in a learning state, and output a prompt message when the probability value is lower than a threshold.

Benefits of technology

It improves the accuracy of detecting whether children are studying, reduces misjudgments, and enhances the effectiveness of supervision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115050054B_ABST
    Figure CN115050054B_ABST
Patent Text Reader

Abstract

This application discloses a prompting method, apparatus, electronic device, and readable storage medium, belonging to the field of artificial intelligence technology. The method includes: acquiring a first image captured by a first camera of the electronic device and a second image captured by a second camera of the electronic device, wherein the first and second cameras capture the same target scene but at different camera angles; acquiring stereoscopic feature information in the target scene based on first feature information of the first image and second feature information of the second image; acquiring a probability value of a target object in the target scene being in a learning state based on the stereoscopic feature information; and outputting prompting information when the probability value is less than a target threshold.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, and specifically relates to a prompting method, device, electronic device, and readable storage medium. Background Technology

[0002] With societal development, parents are paying increasing attention to their children's healthy growth, and how to effectively help children develop good study habits has become an urgent need for parents.

[0003] In existing technology, electronic devices have incorporated a feature to help parents monitor their children's learning. This feature uses the device's camera to capture images in real time and detect whether the object in the image (i.e., the child) is studying. Specifically, if the image captured by the camera includes a photograph of the child, it is also considered that the child is studying.

[0004] It is evident that current methods for detecting whether a child is studying have low accuracy. Summary of the Invention

[0005] The purpose of this application is to provide a prompting method that can solve the problem of low accuracy in existing methods for detecting whether a child is studying.

[0006] In a first aspect, embodiments of this application provide a prompting method, which includes: acquiring a first image captured by a first camera of an electronic device and a second image captured by a second camera of the electronic device, wherein the first camera and the second camera capture the same target scene but at different camera angles; acquiring stereoscopic feature information in the target scene based on first feature information of the first image and second feature information of the second image; acquiring a probability value of a target object in the target scene being in a learning state based on the stereoscopic feature information; and outputting prompting information when the probability value is less than a target threshold.

[0007] Secondly, embodiments of this application provide a prompting device, comprising: a first acquisition module, configured to acquire a first image captured by a first camera of an electronic device and a second image captured by a second camera of the electronic device, wherein the target scene captured by the first camera and the second camera is the same but the camera angles are different; a second acquisition module, configured to acquire stereoscopic feature information in the target scene based on first feature information of the first image and second feature information of the second image; a third acquisition module, configured to acquire a probability value of a target object in the target scene being in a learning state based on the stereoscopic feature information; and an output module, configured to output prompting information when the probability value is less than a target threshold.

[0008] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions, when executed by the processor, implement the steps of the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0010] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0011] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0012] In the embodiments of this application, the electronic device includes a first camera and a second camera. The two cameras can capture images of the target scene from different shooting angles, resulting in a parallax between the first image captured by the first camera and the second image captured by the second camera. Based on this parallax, stereoscopic feature information in the captured target scene can be determined. Furthermore, based on the stereoscopic feature information, the probability value of a target object in the target scene being a stereoscopic object can be obtained. Correspondingly, when the target object is a stereoscopic object, it is assumed that the target object is learning, i.e., in a learning state. Therefore, when the probability value of the target object being in a learning state in the target scene is low, a prompt message is output. It can be seen that, based on the embodiments of this application, the phenomenon that the target object is not a stereoscopic object can be detected more accurately, i.e., whether the child is actually learning can be detected more accurately, thereby improving the accuracy of detecting whether a child is learning. Attached Figure Description

[0013] Figure 1 This is a flowchart of the prompting method according to an embodiment of this application;

[0014] Figures 2 to 6 This is a schematic diagram illustrating the prompting method of an embodiment of this application;

[0015] Figure 7 This is a block diagram of a prompting device according to an embodiment of this application;

[0016] Figure 8 This is one of the hardware structure diagrams of the electronic device according to an embodiment of this application;

[0017] Figure 9 This is the second schematic diagram of the hardware structure of the electronic device according to an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0019] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0020] The prompting method provided in this application embodiment can be executed by the prompting device provided in this application embodiment, or by an electronic device integrating the prompting device, wherein the prompting device can be implemented in hardware or software.

[0021] The prompting method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0022] Figure 1 A flowchart of a prompting method according to an embodiment of this application is shown, exemplified by the method being applied to an electronic device, including:

[0023] Step 110: Acquire the first image captured by the first camera of the electronic device and the second image captured by the second camera of the electronic device. The target scene captured by the first camera and the second camera is the same, but the camera angles are different.

[0024] In this embodiment, the electronic device includes two cameras. Based on the installation positions of the two cameras, they can capture images of the same scene, but at different shooting angles. This results in parallax between the two captured images.

[0025] For example, in a scenario where a child is studying at a desk, an electronic device is placed in front of the child so that two cameras can capture images from different angles.

[0026] This embodiment uses the target scenario as an example for illustration.

[0027] Step 120: Obtain the stereoscopic feature information in the target scene based on the first feature information of the first image and the second feature information of the second image.

[0028] In this step, based on the acquired first image, first feature information of the first image is extracted, and based on the acquired second image, second feature information of the second image is extracted.

[0029] Furthermore, based on the first feature information and the second feature information, the parallax between the two images is obtained, thereby enabling the identification of stereoscopic feature information in the target scene.

[0030] Step 130: Based on the stereo feature information, obtain the probability value of the target object in the target scene being in the learning state.

[0031] In this step, the probability value is used to describe the likelihood that the target object in the target scene is in a learning state. The higher the probability value, the greater the likelihood that the target object is in a learning state.

[0032] Optionally, by combining the feature information in the target scene, the target object in the target scene can be found. Then, by combining the obtained stereo feature information, the probability value that the target object is a stereo object can be obtained, which is used as the probability value obtained in this step.

[0033] For example, in a scenario where children's learning is being monitored, the target is the children being monitored.

[0034] Step 140: If the probability value is less than the target threshold, output a prompt message.

[0035] In this step, a target threshold is set, such as 0.5.

[0036] When the probability value is less than the target threshold, it indicates that the target object in the target scene is less likely to be in a learning state, and a prompt message is output.

[0037] For example, in scenarios where children's learning is being monitored, the prompts are displayed to the supervisors, such as parents.

[0038] For reference, if the probability value is less than the target threshold, it could be that the target object is the object in the photo; or that the target object in the target scene has left; etc.

[0039] Optionally, when the probability value is greater than or equal to the target threshold, it indicates that the target object in the target scene is more likely to be in a learning state.

[0040] For example, in a scenario where a child is being monitored for learning, if the probability value is greater than or equal to the target threshold, it indicates that the child is studying at the table, and the steps in this embodiment are repeated to continue monitoring.

[0041] In the embodiments of this application, the electronic device includes a first camera and a second camera. The two cameras can capture images of the target scene from different shooting angles, resulting in a parallax between the first image captured by the first camera and the second image captured by the second camera. Based on this parallax, stereoscopic feature information in the captured target scene can be determined. Furthermore, based on the stereoscopic feature information, the probability value of a target object in the target scene being a stereoscopic object can be obtained. Correspondingly, when the target object is a stereoscopic object, it is assumed that the target object is learning, i.e., in a learning state. Therefore, when the probability value of the target object being in a learning state in the target scene is low, a prompt message is output. It can be seen that, based on the embodiments of this application, the phenomenon that the target object is not a stereoscopic object can be detected more accurately, i.e., whether the child is actually learning can be detected more accurately, thereby improving the accuracy of detecting whether a child is learning.

[0042] In another embodiment of the prompting method of this application, the electronic device includes a first screen and a second screen, a first camera is located on the first screen, a second camera is located on the second screen, and an angle may be formed between the first screen and the second screen.

[0043] For example, electronic devices include foldable screens, which, when folded, can be folded to form a first screen and a second screen, with an angle between the two screens.

[0044] Optionally, the angle between the first screen and the second screen ranges from 0° to 180°.

[0045] Optionally, to ensure that the first and second cameras can capture images based on the target scene, the angle between the first and second screens is in the range of [60°, 150°].

[0046] For example, see Figure 2 While child 201 is studying, the screen of an electronic device is unfolded and placed in front of child 201, so that the angle X between the two screens reaches a certain value. Camera a (i.e., the first camera) captures the first image 202, and camera b (i.e., the second camera) captures the second image 203. Furthermore, since camera a and camera b are on the same horizontal line, there is a horizontal parallax d between the first image 202 and the second image 203.

[0047] As can be seen, in this embodiment, an electronic device with a foldable screen is provided so as to realize the dual-camera layout of this application by utilizing the structural features of the foldable screen.

[0048] or,

[0049] The electronic device includes a first sub-device and a second sub-device, the first sub-device including a first camera and the second sub-device including a second camera.

[0050] For example, an electronic device includes two camera components that are structurally independent, meaning they can be placed in different locations. Each camera component houses a single camera. Furthermore, both camera components are data-connected to a central control unit. In practical applications, the user can place the two camera components at different locations, and the images captured by both cameras will be transmitted to the control unit.

[0051] Application scenarios include situations where a child is studying, with two camera components placed to the left and right of the child respectively. One camera component captures the first image, and the other captures the second image, with parallax between the first and second images.

[0052] As can be seen, in this embodiment, an electronic device is provided that enables the dual-camera layout of this application to be realized.

[0053] In the flow of the prompting method according to another embodiment of this application, step 140 includes:

[0054] Sub-step A1: If the probability value is less than the target threshold, obtain the target time information.

[0055] Optionally, the target time information can be the current time information.

[0056] Sub-step A2: If the target time information does not match the preset time information, output a prompt message.

[0057] Optionally, the preset time information can be the end time information of a task.

[0058] For example, based on presets in electronic devices, parents can output the end time of learning as preset time information.

[0059] Optionally, mismatches include situations where the target time information is earlier than the preset time information.

[0060] For example, if the current time is earlier than the preset end time for learning, and the system detects that the child is not in a learning state, a prompt message will be output.

[0061] Optionally, if the target time information matches the preset time information, it indicates that a task has been completed, and the detection is stopped.

[0062] For example, the target time information is equal to or later than the preset time information.

[0063] In this embodiment, a time dimension is added. When the probability of the target object in the target scene being in a learning state is low, the time matching result is used to determine whether the end time of a task has been reached. Thus, prompt information is output only before the end time is reached, thereby improving the accuracy of the prompt.

[0064] In the flow of the prompting method according to another embodiment of this application, step 120 includes:

[0065] Sub-step B1: Obtain the disparity feature information between the first image and the second image based on the first feature information of the first image and the second feature information of the second image.

[0066] Optionally, the disparity feature information includes feature information where disparity exists in the target scene.

[0067] Sub-step B2: Based on the first feature information of the first image and the second feature information of the second image, obtain the positional deviation information between the first image and the second image.

[0068] Optionally, positional deviation information is used to describe the positional deviation between the two images.

[0069] Sub-step B3: Obtain stereo feature information in the target scene based on disparity feature information and positional deviation information.

[0070] In this step, the disparity feature information can be multiplied with the positional deviation information, which is equivalent to assigning a certain weight to the disparity feature information, thus filtering out the truly effective disparity feature information.

[0071] Furthermore, by performing element-by-element addition and fusion filtering, three-dimensional feature information is obtained.

[0072] This embodiment provides a method for acquiring stereo feature information. Based on the parallax features present in the target scene and combined with the positional deviation between two images, stereo feature information is ultimately obtained. It is evident that, considering various factors, the acquired stereo feature information is relatively accurate and has practical reference value.

[0073] In the flow of the prompting method according to another embodiment of this application, step B1 includes:

[0074] Sub-step C1: Obtain the first fused feature information based on the first feature information of the first image and the second feature information of the second image.

[0075] Optionally, feature extraction and feature fusion can be performed using a feature extraction module.

[0076] For example, see Figure 3 The feature extraction module uses MobileNetV3 as its main structure and outputs two sets of different semantic layers of MobileNetV3: {C11, C12, C13, C14} and {C21, C22, C23, C24}. Figure 3 In the image, the left and right images represent the first and second images, respectively.

[0077] See Figure 3 The MobileNetV3 module above contains four convolutional modules {C11, C12, C13, C14}, while the MobileNetV3 module below contains convolutional modules {C21, C22, C23, C24}. C11 contains three sub-modules with a 3×3 kernel size, producing 16, 24, and 24 output channels respectively; C12 contains three sub-modules with a 5×5 kernel size, producing 40, 40, and 40 output channels respectively; C13 contains six sub-modules with a 3×3 kernel size, producing 80, 80, 80, 80, 112, and 112 output channels respectively; and C14 contains three sub-modules with a 5×5 kernel size, producing 160, 160, and 160 output channels respectively. Each sub-module is constructed using three convolutional layers and one skip layer with a 1×1 kernel size. {C21, C22, C23, C24} adopts the same design as {C11, C12, C13, C14}.

[0078] Among them, see Figure 3 The output features of the convolutional modules {C11, C12, C13, C14} in the feature extraction network above are {C 11 C 12 C 13 C 14 The output features of the convolutional modules {C21, C22, C23, C24} in the feature extraction network below are {C... 21 C 22 C 23 C 24}

[0079] Correspondingly, {C 11 C 12 C 13 C 14} and {C 21 C 22C 23 C 24} are used to represent the first feature information and the second feature information, respectively.

[0080] Furthermore, the output features of the two MobileNetV3 semantic layers from different semantic layers are stacked.

[0081] Where, C4'=H([C 14 C 24 ]);

[0082] C3'=H([C 13 C 23 ,U(C4')]);

[0083] C2'=H([C 12 C 22 ,U(C3')]);

[0084] C1'=H([C 11 C 21 ,U(C2')]).

[0085] Where H represents the feature stacking operation, implemented through the concat operation, and U represents the doubling upsampling operation; {C1', C2', C3', C4'} represent the features after the fusion of the upper and lower feature extraction networks, respectively, and their feature map sizes are the same as {C 11 C 12 C 13 C 14}、{C 21 C 22 C 23 C 24 All are consistent; the number of output channels of C4' is C. 14 The number of output channels of C3' is twice that of C. 13 The number of output channels of C2' is 4 times that of C. 12 The number of output channels of C1' is 6 times that of C. 11 8 times.

[0086] Sub-step C2: Based on the first fused feature information, obtain the first feature map dimension information, which includes the maximum disparity value, image height value, image width value, and feature map channel number value.

[0087] Optionally, the first feature map dimension information can be output using the Cost Volume module.

[0088] For example, see Figure 4 , will merge {C 11 C 12 C 13 C14} and {C 21 C 22 C 23 C 24 The inputs are all fed into the CostVolume module. Through the reshape operation, the output feature map dimensions are reshaped (i.e., the first feature map dimension information) to {d, h, w, c, d}.

[0089] Where d represents the maximum disparity (e.g., 160), h represents the height of the input image, w represents the width of the input image, and c represents the number of channels in the fused feature map (e.g., 1280).

[0090] In addition, d is included in the feature map dimension. Its purpose is to enable the subsequent feature encoder to extract disparity features. The maximum possible disparity size is set, and the effective disparity feature information is then filtered out by the channel attention module.

[0091] Sub-step C3: Based on the dimensional information of the first feature map, obtain the disparity feature information between the first image and the second image.

[0092] Optionally, the Cost Volume module can be used to output disparity feature information.

[0093] For example, see Figure 4 The feature map obtained after the reshape operation is input into the feature encoder, which consists of three consecutive 3-dimensional (3D) convolution operations. The first 3D convolution has d input channels and d / 2 output channels. The second 3D convolution has d / 2 input and d / 4 output channels. The stride of each convolution is {1, 1, 1}, the kernel size is {3, 3, 3}, and the padding size is {1, 1, 1}. 3D convolutions can extract higher-dimensional spatial feature information; in this case, disparity features are extracted.

[0094] In the embodiments of this application, the feature information can also be represented in the form of channels. That is, the feature information includes feature channels.

[0095] In this embodiment, a method for obtaining disparity feature information is provided, which involves multiple steps such as feature extraction, feature fusion, feature map dimension acquisition, and feature encoding to obtain disparity feature information.

[0096] In the flow of the prompting method according to another embodiment of this application, step B2 includes:

[0097] Sub-step D1: Divide the first image and the second image into N small blocks respectively, where N is a positive integer.

[0098] In this embodiment, there will be a certain positional deviation at the same location point in the first image and the second image. Therefore, the purpose of this embodiment is to obtain the positional deviation information between the two images.

[0099] In this step, the two images are first divided into multiple small blocks. The more blocks are divided, the higher the matching accuracy will be, but the amount of computation will increase accordingly. Therefore, the number of blocks can be appropriately determined.

[0100] The two images are divided into the same number of small blocks to facilitate subsequent difference calculations.

[0101] Sub-step D2: Obtain the feature information corresponding to each small block.

[0102] In this step, feature extraction is performed on the sub-image corresponding to each small block. Here, global average pooling is used to obtain a single feature value.

[0103] Sub-step D3: Based on the feature information corresponding to the small blocks in the first image and the feature information corresponding to the small blocks in the second image, obtain the positional deviation information between the first image and the second image.

[0104] The feature values ​​corresponding to all small blocks in the first image constitute the feature vector of the first image, and the feature values ​​corresponding to all small blocks in the second image constitute the feature vector of the second image. The difference between the feature vectors of the two images is calculated to obtain the distance vector between the two images.

[0105] In this embodiment, the positional deviation information is represented by a distance vector.

[0106] For example, see Figure 5 The distance vector is obtained from the feature vectors based on the two images. The obtained distance vector is input into the fully connected layer, and its output is a feature vector of {1, 160}. Then, it is activated by the Sigmoid function to obtain each value of the vector between 0 and 1.

[0107] The final feature vector is actually a mask used to filter out the useful parallax channel portion in the Cost Volume module.

[0108] Optionally, the channel attention module can be used to output positional deviation information.

[0109] In this embodiment, a method for obtaining positional deviation information is provided, which obtains positional deviation information by calculating the difference between the features of two images.

[0110] In another embodiment of the prompting method of this application, if the two cameras are at the same level, there is only horizontal parallax between the two images, and the positional deviation information, i.e. the magnitude of the horizontal parallax, can be obtained directly by pattern matching.

[0111] For example, iterate through the distance of each pixel, and then, at the current distance, calculate the norm distance between the first and second images starting from that distance. The pixel distance that is minimized is the magnitude of the horizontal parallax.

[0112] In the flow of the prompting method according to another embodiment of this application, step 130 includes:

[0113] Sub-step E1: Obtain the dimension information of the second feature map based on the 3D feature information.

[0114] In this step, the fused and filtered 3D feature information is input into the Cost Volume module. Through the reshape operation, the dimensional information of the second feature map is obtained, which is used to obtain the probability value later.

[0115] For example, the feature map dimension reshape (i.e., the second feature map dimension information) is {1, 1296}.

[0116] Sub-step E2: Based on the dimensional information of the second feature map, obtain the probability value of the target object in the target scene being in the learning state.

[0117] Alternatively, in conjunction with the two embodiments described above, see [link to relevant documentation]. Figure 6 The obtained distance vector (i.e. the angular feature vector in the figure) is multiplied by the output of the Cost Volume module and input into the fully connected layer, from which the probability value can be output.

[0118] The disparity feature information output by the Cost Volume module is multiplied by the distance vector to further filter the disparity feature information, thereby filtering out the effective disparity feature information for use in probability value calculation.

[0119] Optionally, the probability value can be output using the channel attention module.

[0120] For example, the channel attention module is built using a fully connected layer. The fully connected layer has 1296 input channels and 1 output channel. Then, a sigmoid layer is added after the fully connected layer to obtain the probability value of the classification target.

[0121] In this context, the target object in the target scene being in a learning state is considered as a classification target, in order to obtain the probability value of that classification target.

[0122] In this embodiment, a method for outputting a probability value is provided, in which a probability value is output by the channel attention module based on stereo feature information.

[0123] In the flow of the prompting method according to another embodiment of this application, after step 140, the method further includes:

[0124] Step F1: Send a prompt message to the target electronic device via data connection.

[0125] Optionally, the target electronic device and the electronic device of this application are connected via Bluetooth data connection.

[0126] Correspondingly, a notification message is sent to the target electronic device via Bluetooth data.

[0127] For example, the target electronic device is a wearable device such as a watch.

[0128] Optionally, the target electronic device and the electronic device of this application are connected via a mobile network data connection.

[0129] Correspondingly, the identification information of the target electronic device, such as its phone number, can be preset, thereby sending a prompt message to the target electronic device via mobile network data.

[0130] For example, the target electronic device is a mobile phone or similar device.

[0131] In this embodiment, a prompt message can be sent to the associated electronic device so that the other party can receive the corresponding prompt. This is applicable to scenarios where parents supervise their children's learning, etc.

[0132] In the embodiments of this application, users can manually turn the detection function provided by this application on and off; or they can automatically turn the detection function provided by this application on and off according to the user's time settings.

[0133] In the embodiments of this application, depending on the electronic device, the size of the module needs to meet certain requirements. Therefore, the number of module parameters can be reduced by using group convolution or depthwise separable convolution.

[0134] In summary, this application combines the parallax generated by images acquired from dual cameras to capture the target object being learned in a scene, achieving the effect of supervised learning. In application, images from multiple perspectives are acquired using dual cameras, multi-dimensional features are extracted and fused through a convolutional neural network, and then the probability value of the target object being in a learning state is calculated based on the convolutional neural network. Furthermore, a dynamic parallax dual-camera algorithm is constructed using an attention mechanism to overcome the problem of decreased monitoring performance caused by different acquisition angles. This application, when applied to learning and other detection scenarios, can effectively prevent cheating and achieve efficient and intelligent detection of children's learning.

[0135] The prompting method provided in this application can be executed by a prompting device. This application uses the example of a prompting device executing the prompting method to illustrate the prompting device provided in this application.

[0136] Figure 7 A block diagram of a prompting device according to another embodiment of this application is shown, the device comprising:

[0137] The first acquisition module 10 is used to acquire a first image captured by the first camera of the electronic device and a second image captured by the second camera of the electronic device. The target scene captured by the first camera and the second camera is the same, but the camera angles are different.

[0138] The second acquisition module 20 is used to acquire stereoscopic feature information in the target scene based on the first feature information of the first image and the second feature information of the second image;

[0139] The third acquisition module 30 is used to acquire the probability value of the target object in the target scene being in the learning state based on the stereo feature information;

[0140] Output module 40 is used to output a prompt message when the probability value is less than the target threshold.

[0141] In the embodiments of this application, the electronic device includes a first camera and a second camera. The two cameras can capture images of the target scene from different shooting angles, resulting in a parallax between the first image captured by the first camera and the second image captured by the second camera. Based on this parallax, stereoscopic feature information in the captured target scene can be determined. Furthermore, based on the stereoscopic feature information, the probability value of a target object in the target scene being a stereoscopic object can be obtained. Correspondingly, when the target object is a stereoscopic object, it is assumed that the target object is learning, i.e., in a learning state. Therefore, when the probability value of the target object being in a learning state in the target scene is low, a prompt message is output. It can be seen that, based on the embodiments of this application, the phenomenon that the target object is not a stereoscopic object can be detected more accurately, i.e., whether the child is actually learning can be detected more accurately, thereby improving the accuracy of detecting whether a child is learning.

[0142] Optionally, the electronic device includes a first screen and a second screen, with a first camera located on the first screen and a second camera located on the second screen, and an angle may be formed between the first screen and the second screen; or, the electronic device includes a first sub-device and a second sub-device, with the first sub-device including a first camera and the second sub-device including a second camera.

[0143] Optionally, the output module 40 includes:

[0144] The first acquisition unit is used to acquire target time information when the probability value is less than the target threshold.

[0145] The output unit is used to output a prompt message when the target time information does not match the preset time information.

[0146] Optionally, the second acquisition module 20 includes:

[0147] The second acquisition unit is used to acquire disparity feature information between the first image and the second image based on the first feature information of the first image and the second feature information of the second image;

[0148] The third acquisition unit is used to acquire positional deviation information between the first image and the second image based on the first feature information of the first image and the second feature information of the second image;

[0149] The fourth acquisition unit is used to acquire stereo feature information in the target scene based on disparity feature information and position deviation information.

[0150] Optionally, the second acquisition unit includes:

[0151] The first sub-acquisition unit is used to acquire the fused first fused feature information based on the first feature information of the first image and the second feature information of the second image;

[0152] The second sub-acquisition unit is used to acquire the first feature map dimension information based on the first fusion feature information. The first feature map dimension information includes the maximum disparity value, the image height value, the image width value, and the number of feature map channels.

[0153] The third sub-acquisition unit is used to acquire disparity feature information between the first image and the second image based on the dimension information of the first feature map.

[0154] Optionally, the third acquisition unit includes:

[0155] The splitting sub-unit is used to split the first image and the second image into N small blocks respectively, where N is a positive integer;

[0156] The fourth sub-acquisition unit is used to acquire the feature information corresponding to each small block respectively;

[0157] The fifth sub-acquisition unit is used to acquire positional deviation information between the first image and the second image based on the feature information corresponding to the small blocks in the first image and the feature information corresponding to the small blocks in the second image.

[0158] Optionally, the third acquisition module 30 includes:

[0159] The fifth acquisition unit is used to acquire the dimension information of the second feature map based on the stereo feature information;

[0160] The sixth acquisition unit is used to obtain the probability value of the target object in the target scene being in the learning state based on the dimensional information of the second feature map.

[0161] The prompting device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the scope of the device.

[0162] The prompting device in this application embodiment can be a device with an action system. The action system can be an Android action system, an iOS action system, or other possible action systems; this application embodiment does not specifically limit it.

[0163] The prompting device provided in this application embodiment can implement the various processes implemented in the above method embodiments, and will not be described again here to avoid repetition.

[0164] Optionally, such as Figure 8 As shown, this application embodiment also provides an electronic device 100, including a processor 101, a memory 102, and a program or instructions stored in the memory 102 and executable on the processor 101. When the program or instructions are executed by the processor 101, they implement the various steps of any of the above-described prompting method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0165] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0166] Figure 9 A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0167] The electronic device 1000 includes, but is not limited to, components such as: radio frequency unit 1001, network module 1002, audio output unit 1003, input unit 1004, sensor 1005, display unit 1006, user input unit 1007, interface unit 1008, memory 1009, and processor 1010.

[0168] Those skilled in the art will understand that the electronic device 1000 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1010 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 9 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0169] The processor 1010 is configured to acquire a first image captured by a first camera of the electronic device and a second image captured by a second camera of the electronic device, wherein the target scene captured by the first camera and the second camera is the same but the camera angles are different; acquire stereoscopic feature information in the target scene based on the first feature information of the first image and the second feature information of the second image; acquire the probability value of the target object in the target scene being in a learning state based on the stereoscopic feature information; and output a prompt message if the probability value is less than a target threshold.

[0170] In the embodiments of this application, the electronic device includes a first camera and a second camera. The two cameras can capture images of the target scene from different shooting angles, resulting in a parallax between the first image captured by the first camera and the second image captured by the second camera. Based on this parallax, stereoscopic feature information in the captured target scene can be determined. Furthermore, based on the stereoscopic feature information, the probability value of a target object in the target scene being a stereoscopic object can be obtained. Correspondingly, when the target object is a stereoscopic object, it is assumed that the target object is learning, i.e., in a learning state. Therefore, when the probability value of the target object being in a learning state in the target scene is low, a prompt message is output. It can be seen that, based on the embodiments of this application, the phenomenon that the target object is not a stereoscopic object can be detected more accurately, i.e., whether the child is actually learning can be detected more accurately, thereby improving the accuracy of detecting whether a child is learning.

[0171] Optionally, the electronic device includes a first screen and a second screen, with the first camera located on the first screen and the second camera located on the second screen, and the first screen and the second screen may form an angle between them; or, the electronic device includes a first sub-device and a second sub-device, with the first sub-device including the first camera and the second sub-device including the second camera.

[0172] Optionally, the processor 1010 is further configured to acquire target time information when the probability value is less than a target threshold, and output the prompt information when the target time information does not match the preset time information.

[0173] Optionally, the processor 1010 is further configured to obtain disparity feature information between the first image and the second image based on the first feature information of the first image and the second feature information of the second image; obtain positional deviation information between the first image and the second image based on the first feature information of the first image and the second feature information of the second image; and obtain stereoscopic feature information in the target scene based on the disparity feature information and the positional deviation information.

[0174] Optionally, the processor 1010 is further configured to: obtain fused first fused feature information based on the first feature information of the first image and the second feature information of the second image; obtain first feature map dimension information based on the first fused feature information, wherein the first feature map dimension information includes the maximum disparity value, image height value, image width value, and feature map channel number value; and obtain disparity feature information between the first image and the second image based on the first feature map dimension information.

[0175] Optionally, the processor 1010 is further configured to divide the first image and the second image into N small blocks, where N is a positive integer; obtain feature information corresponding to each of the small blocks; and obtain positional deviation information between the first image and the second image based on the feature information corresponding to the small blocks of the first image and the feature information corresponding to the small blocks of the second image.

[0176] Optionally, the processor 1010 is further configured to obtain second feature map dimension information based on the stereo feature information; and to obtain the probability value of the target object in the target scene being in a learning state based on the second feature map dimension information.

[0177] In summary, this application combines the parallax generated by images acquired from dual cameras to capture the target object being learned in a scene, achieving the effect of supervised learning. In application, images from multiple perspectives are acquired using dual cameras, multi-dimensional features are extracted and fused through a convolutional neural network, and then the probability value of the target object being in a learning state is calculated based on the convolutional neural network. Furthermore, a dynamic parallax dual-camera algorithm is constructed using an attention mechanism to overcome the problem of decreased monitoring performance caused by different acquisition angles. This application, when applied to learning and other detection scenarios, can effectively prevent cheating and achieve efficient and intelligent detection of children's learning.

[0178] It should be understood that, in this embodiment, the input unit 1004 may include a graphics processing unit (GPU) 10041 and a microphone 10042. The GPU 10041 processes image data of still images or video images obtained by an image capture device (such as a camera) in video image capture mode or image capture mode. The display unit 1006 may include a display panel 10061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 1007 includes a touch panel 10071 and at least one of other input devices 10072. The touch panel 10071 is also called a touch screen. The touch panel 10071 may include a touch detection device and a touch controller. Other input devices 10072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here. The memory 1009 can be used to store software programs and various data, including but not limited to applications and motion systems. Processor 1010 may integrate an application processor and a modem processor. The application processor mainly handles the action system, user page, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 1010.

[0179] The memory 1009 can be used to store software programs and various data. The memory 1009 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1009 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1009 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0180] The processor 1010 may include one or more processing units; optionally, the processor 1010 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into the processor 1010.

[0181] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described prompting method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0182] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0183] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-mentioned method embodiments and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0184] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0185] This application provides a computer program product that is stored in a storage medium and executed by at least one processor to implement the various processes described in the above-described method embodiments, achieving the same technical effects. To avoid repetition, further details are omitted here.

[0186] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0187] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0188] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A prompting method characterized by comprising: The method comprises: obtaining a first image collected by a first camera of an electronic device and a second image collected by a second camera of the electronic device, the target scene photographed by the first camera and the second camera being the same, and the camera angles being different; obtaining stereo feature information in the target scene according to first feature information of the first image and second feature information of the second image, wherein the stereo feature information is obtained by multiplying parallax feature information and position deviation information, element-by-element addition and fusion screening; the parallax feature information comprises feature information with parallax in the target scene, and the position deviation information is used for describing the position deviation between the first image and the second image; and the position deviation information comprises a distance vector between the first image and the second image; obtaining second feature map dimension information according to the stereo feature information; obtaining a probability value of a target object in the target scene being in a learning state according to the second feature map dimension information; outputting prompt information in a case where the probability value is less than a target threshold.

2. The method of claim 1, wherein, The electronic device comprises a first screen and a second screen, the first camera is located on the first screen, the second camera is located on the second screen, and an included angle can be formed between the first screen and the second screen; or The electronic device comprises a first sub-device and a second sub-device, the first sub-device comprises the first camera, and the second sub-device comprises the second camera.

3. The method of claim 1, wherein, The outputting prompt information in the case where the probability value is less than the target threshold comprises: obtaining target time information in the case where the probability value is less than the target threshold; outputting the prompt information in a case where the target time information does not match preset time information.

4. The method of claim 1, wherein, The obtaining stereo feature information in the target scene according to the first feature information of the first image and the second feature information of the second image comprises: obtaining parallax feature information between the first image and the second image according to the first feature information of the first image and the second feature information of the second image; obtaining position deviation information between the first image and the second image according to the first feature information of the first image and the second feature information of the second image; obtaining stereo feature information in the target scene according to the parallax feature information and the position deviation information.

5. The method of claim 4, wherein, The obtaining parallax feature information between the first image and the second image according to the first feature information of the first image and the second feature information of the second image comprises: obtaining first fusion feature information after fusion according to the first feature information of the first image and the second feature information of the second image; obtaining first feature map dimension information according to the first fusion feature information, the first feature map dimension information comprising a maximum parallax size value, an image height value, an image width value and a feature map channel number value; obtaining parallax feature information between the first image and the second image according to the first feature map dimension information.

6. The method of claim 4, wherein, The first feature information of the first image and the second feature information of the second image are used to obtain position deviation information between the first image and the second image. The first image and the second image are respectively split into N small blocks, where N is a positive integer. Feature information corresponding to each small block is obtained. The feature information corresponding to the small blocks of the first image and the feature information corresponding to the small blocks of the second image are used to obtain position deviation information between the first image and the second image.

7. A prompting device, characterized by The device comprises: A first obtaining module is configured to obtain a first image captured by a first camera of an electronic device and a second image captured by a second camera of the electronic device, wherein the first camera and the second camera capture the same target scene but have different camera angles. A second obtaining module is configured to obtain stereoscopic feature information in the target scene according to first feature information of the first image and second feature information of the second image, wherein the stereoscopic feature information is obtained by multiplying, element-wise adding, and fusing and screening disparity feature information and position deviation information; the disparity feature information includes feature information with disparity in the target scene, and the position deviation information is used to describe a position deviation between the first image and the second image; and the position deviation information includes a distance vector between the first image and the second image. A third obtaining module is configured to obtain a probability value of a target object in the target scene being in a learning state according to the stereoscopic feature information. An output module is configured to output prompt information when the probability value is less than a target threshold. The third obtaining module comprises: A fifth obtaining unit is configured to obtain second feature map dimension information according to the stereoscopic feature information. A sixth obtaining unit is configured to obtain a probability value of a target object in the target scene being in a learning state according to the second feature map dimension information.

8. The apparatus of claim 7, wherein, The electronic device comprises a first screen and a second screen, the first camera is located on the first screen, the second camera is located on the second screen, and an included angle can be formed between the first screen and the second screen; or The electronic device comprises a first sub-device and a second sub-device, the first sub-device comprises the first camera, and the second sub-device comprises the second camera.

9. The apparatus of claim 7, wherein, The output module comprises: A first obtaining unit is configured to obtain target time information when the probability value is less than a target threshold. An output unit is configured to output the prompt information when the target time information does not match preset time information.

10. The apparatus of claim 7, wherein, The second obtaining module comprises: A second obtaining unit is configured to obtain disparity feature information between the first image and the second image according to first feature information of the first image and second feature information of the second image. A third obtaining unit is configured to obtain position deviation information between the first image and the second image according to first feature information of the first image and second feature information of the second image. A fourth obtaining unit, configured to obtain stereo feature information in the target scene according to the parallax feature information and the position deviation information.

11. The apparatus of claim 10, wherein, The second obtaining unit comprises: A first sub-obtaining unit, configured to obtain first fused feature information according to first feature information of the first image and second feature information of the second image; A second sub-obtaining unit, configured to obtain first feature map dimension information according to the first fused feature information, the first feature map dimension information comprising a maximum parallax value, an image height value, an image width value and a feature map channel number value; A third sub-obtaining unit, configured to obtain parallax feature information between the first image and the second image according to the first feature map dimension information.

12. The apparatus of claim 10, wherein, The third obtaining unit comprises: A splitting sub-unit, configured to split the first image and the second image into N small blocks respectively, N being a positive integer; A fourth sub-obtaining unit, configured to obtain feature information corresponding to each small block respectively; A fifth sub-obtaining unit, configured to obtain position deviation information between the first image and the second image based on the feature information corresponding to the small blocks of the first image and the feature information corresponding to the small blocks of the second image.

13. An electronic device, comprising: A processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the steps of the prompting method according to any one of claims 1-6.

14. A readable storage medium, characterized by, The readable storage medium stores programs or instructions, the programs or instructions being executed by the processor to implement the steps of the prompting method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Image stereo matching method based on PSMNet optimization

    CN112150521A

  • Circular cooler trolley wheel detection method, device and equipment and medium

    CN112393617A

  • Image information processing method and device and electronic equipment

    CN113301320A

  • Remote teaching assistance method and device, equipment and storage medium

    CN113468930A

  • Medical equipment machine vision image processing method and computer readable storage medium

    CN114331855A