Abnormal monitoring method and device, mobile terminal and computer readable storage medium
By combining the first camera and the second camera of the mobile terminal to detect abnormal changes and events, the problems of high cost, poor accuracy and dependence on network conditions in the existing technology are solved, and low-cost, efficient abnormal monitoring and early warning of mobile terminals are achieved.
Patent Information
- Application Number
- CN202410346504.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-25
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies require the installation of additional network cameras for user environment monitoring, which is costly, has low pixels, and depends on network conditions, resulting in poor accuracy and delays in abnormal monitoring. In addition, mobile terminal applications lack effective abnormal monitoring solutions.
The first camera of the mobile terminal is used to capture video images with a wide viewing angle, and abnormal changes in the picture are detected through multi-frame images. The second camera is combined to capture high-pixel video images for abnormal event detection to achieve accurate early warning.
It achieves low-cost, fast and accurate anomaly monitoring and early warning to ensure user safety, without the need for additional hardware installation and independent of network conditions.
Smart Images

Figure CN120711142A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of mobile terminal application technology, and more specifically, to an abnormality monitoring method, device, mobile terminal and computer-readable storage medium. Background Art
[0002] Currently, most methods for monitoring abnormalities in a user's environment involve installing independent network cameras indoors. These cameras capture the user's surroundings and transmit the captured footage to a cloud server, which then performs abnormality monitoring based on the captured footage. This method has the following drawbacks: ① The cost of installing an additional network camera is high; ② The pixel count of network cameras is low, resulting in poor accuracy in abnormality monitoring; and ③ The transmission from the network camera to the cloud server is limited by the network conditions between the network camera and the cloud server, which can result in delays in abnormality monitoring when the network is poor. Summary of the Invention
[0003] The embodiments of the present application provide an abnormality monitoring method, device, mobile terminal, and computer-readable storage medium to improve the above-mentioned problems.
[0004] In a first aspect, an embodiment of the present application provides an abnormal monitoring method, which is applied to a mobile terminal, wherein the mobile terminal includes a first camera and a second camera, the viewing angle of the first camera is larger than that of the second camera, and the method includes: acquiring a first video image captured by the first camera and a second video image captured by the second camera; detecting whether an abnormal change occurs in the monitoring screen based on multiple frames of images in the first video image; when an abnormal change occurs in the monitoring screen, detecting an abnormal event based on the second video image; and issuing an early warning message based on the abnormal event.
[0005] In a second aspect, an embodiment of the present application provides an abnormality monitoring device, which is applied to a mobile terminal, wherein the mobile terminal includes a first camera and a second camera, the viewing angle of the first camera is larger than that of the second camera, and the device includes: an image acquisition module, which is used to capture a first video image by the first camera and a second video image by the second camera; an abnormality detection module, which is used to detect whether an abnormal change occurs in the monitoring screen based on multiple frames of images in the first video image; an event detection module, which is used to detect an abnormal event based on the second video image when an abnormal change occurs in the monitoring screen; and an abnormal warning module, which is used to issue a warning message based on the abnormal event.
[0006] In a third aspect, an embodiment of the present application provides a mobile terminal comprising: a first camera, a second camera, a memory and a processor, wherein an application is stored in the memory, and when the processor calls the application, the method provided in the embodiment of the present application is executed.
[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having program code stored thereon, and when a processor calls the program code, the method provided in the embodiment of the present application is executed.
[0008] The abnormality monitoring method, device, mobile terminal, and computer-readable storage medium provided in the embodiments of the present application simultaneously use the first camera of the mobile terminal to capture a first video image and the second camera to capture a second video image. Due to the wider viewing angle of the first camera, the use of multiple frames of the first video image can more quickly detect whether abnormal changes have occurred in the monitoring screen. When abnormal changes in the monitoring screen are detected, the second video image is used to detect abnormal events, which can accurately determine the abnormal event and achieve accurate early warning of abnormal events, ensuring the personal safety of users. This expands the application of mobile terminals, making it possible to implement abnormality monitoring and early warning through mobile terminals, and can ensure the personal safety of users at a low cost.
[0009] In addition, compared with the method of abnormality monitoring based on network cameras and cloud servers, this application has at least the following advantages:
[0010] ① This application uses a single mobile terminal device to achieve abnormal monitoring and early warning, without the need to install a network camera, which can reduce costs;
[0011] ② The camera pixels of mobile terminals are higher than those of network cameras, and abnormal monitoring and abnormal warning are more accurate;
[0012] ③ The data exchanged between the mobile terminal and its own camera is basically unaffected by network conditions, and abnormal monitoring is faster and more robust. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] To more clearly illustrate the technical solutions in the embodiments of this application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of this application, not all embodiments. All other embodiments and drawings obtained by ordinary technicians in this field based on the embodiments of this application without creative work are within the scope of protection of this application.
[0014] Figure 1 A schematic diagram showing an abnormality monitoring system architecture of the related art;
[0015] Figure 2A schematic diagram showing a hardware operating environment of the abnormality monitoring method provided in an embodiment of the present application;
[0016] Figure 3 A schematic diagram of a mobile terminal provided in an embodiment of the present application is shown;
[0017] Figure 4 A schematic diagram of a mobile terminal provided by another embodiment of the present application is shown;
[0018] Figure 5 A schematic diagram showing a monitoring scanning range provided by an exemplary embodiment of the present application is shown;
[0019] Figure 6 A schematic diagram of a mobile terminal provided in another embodiment of the present application is shown;
[0020] Figure 7 A schematic diagram of a mobile terminal provided in yet another embodiment of the present application is shown;
[0021] Figure 8 A schematic diagram of a process flow of an abnormality monitoring method provided by an embodiment of the present application is shown;
[0022] Figure 9 A schematic diagram of human joints provided by an exemplary embodiment of the present application is shown;
[0023] Figure 10 A schematic diagram of the OpenPose network architecture provided in one embodiment of the present application is shown;
[0024] Figure 11 A partial flow chart of an abnormality monitoring method provided by another embodiment of the present application is shown;
[0025] Figure 12 A schematic diagram of a mobile terminal display interface provided by an exemplary embodiment of the present application is shown;
[0026] Figure 13 A schematic diagram showing a mobile terminal display interface provided by another exemplary embodiment of the present application is shown;
[0027] Figure 14 A schematic diagram of a mobile terminal display interface provided by another exemplary embodiment of the present application is shown;
[0028] Figure 15 A schematic diagram of a mobile terminal display interface provided by yet another exemplary embodiment of the present application is shown;
[0029] Figure 16 A schematic diagram of a mobile terminal provided in another embodiment of the present application is shown;
[0030] Figure 17A schematic diagram of a mobile terminal provided in yet another embodiment of the present application is shown;
[0031] Figure 18 A schematic diagram showing a flow chart of an abnormality monitoring method provided by another embodiment of the present application is shown;
[0032] Figure 19 A schematic structural diagram of an abnormality monitoring device provided in an embodiment of the present application is shown;
[0033] Figure 20 A structural diagram of a mobile terminal provided in another embodiment of the present application is shown. DETAILED DESCRIPTION
[0034] In order to enable people skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0035] When users travel or go on business trips, they need to stay in hotels. Since the security measures of various types of hotels vary, users will inevitably worry about their personal safety when sleeping at night, such as strangers breaking in or other dangerous situations.
[0036] In order to protect the personal safety of users, Figure 1 As shown, most current approaches to abnormality monitoring use network cameras and cloud servers. This approach requires network cameras, object storage, cloud servers, and clients, and also involves operations such as video stream storage and transmission, user authentication, and other related operations. This approach involves the following: First, a network camera or fixed camera captures the video. Cameras are typically mounted on walls or rooftops, for example, in locations with wide viewing angles to capture more footage. Due to cost constraints, network cameras typically have low pixel counts. The video stream captured by the camera is then transmitted over the network to an object storage service (OSS) for storage. This storage process typically involves compression and encryption, as well as user access rights verification. The cloud server manages the storage space. Finally, users and administrators can monitor the video surveillance footage through authorization to handle and respond to any dangerous situations captured in the video.
[0037] The inventors have found in long-term research that the above-mentioned abnormality monitoring method has the following defects: ① It requires auxiliary installation and wiring of hardware equipment, which is complicated to operate and costly; ② The pixel of the network camera is low, and the accuracy of the existing abnormality monitoring is insufficient; ③ It requires separate software and hardware support and the hardware is large, which is often troublesome or inconvenient to carry, and requires a certain network environment support. It is limited by the network conditions. If the network conditions are not good, abnormality monitoring may be delayed; ④ The storage of audio and video data of the monitoring system requires TF (Trans-flash Card, currently corrected to Micro SD Card) card or cloud storage space support; ⑤ The monitoring system needs to be manually monitored to quickly deal with dangerous situations.
[0038] In addition, the inventors have also found in their long-term research that the monitoring products currently on the market all use the above-mentioned solution as the basis for their architecture and design, and the form of the entire technical solution is relatively fixed. In other words, the current abnormal monitoring solutions are basically designed based on indoor monitoring cameras, and there are almost no abnormal monitoring solutions based on mobile terminals (such as smart phones), which represents a blank market space. The imaging position, imaging environment, and hardware user requirements of mobile terminals are very different from those of the above-mentioned monitoring systems. For example, the imaging position, imaging environment, and hardware user requirements of foldable screen mobile phones in hover mode are very different from those of the above-mentioned monitoring systems. Therefore, the above-mentioned security monitoring solution is not applicable to mobile terminals.
[0039] Based on this, the inventors have proposed a method for monitoring abnormalities. This method simultaneously uses a first camera on a mobile terminal to capture a first video image, and a second camera to capture a second video image. Multiple frames of the first video image, which has a wider viewing angle, are used to quickly detect abnormal changes in the monitoring image. When an abnormal change in the monitoring image is detected, the second video image is used to further detect the abnormal event, thereby achieving accurate early warning of the abnormal event and ensuring the personal safety of the user. This expands the application of mobile terminals, enabling automatic abnormality monitoring and early warning through mobile terminals, ensuring the personal safety of users at a low cost. This method overcomes the aforementioned shortcomings of abnormality monitoring based on network cameras and cloud servers.
[0040] The anomaly monitoring method provided in this application can be applied to an anomaly monitoring device or a mobile terminal. Specifically, the anomaly monitoring device can be applied to a mobile terminal. Mobile terminals can include, but are not limited to, smartphones, tablets, and laptops. Smartphones can include, but are not limited to, foldable phones, flat-screen phones, and curved-screen phones.
[0041] The hardware / environmental resources of this application are mainly concentrated on mobile terminals. Due to the use of deep model processing, there are certain requirements for processors, memory, and storage space. For example, the system of the mobile terminal can be Android 4.4 and above, the central processing unit (CPU) can be quad-core, the running memory can be greater than or equal to 4GB, and the storage space can be more than 32GB. This application can also run on the IOS system, the IOS version can be IOS9.0 and above, and the recommended standard for supported devices is that the random access memory (RAM) can be greater than 2GB and the CPU can be A9 and above.
[0042] The software and hardware system architecture of the mobile terminal 100 is as follows: Figure 2 As shown, the mobile terminal 100 may include: a user interaction system 110, a monitoring and analysis system 120, and a hardware system 130. The user interaction system 110 is used to interact with the user. The monitoring and analysis system 120 is used to perform anomaly analysis and early warning based on the monitored video images. The hardware system 130 is used to provide the various hardware required for anomaly monitoring.
[0043] The user interaction system 110 may include a monitoring screen display module 111, a voice recognition module 112, and a system settings module 113. The monitoring screen display module 111 displays the monitoring screen in real time for the user to view. The voice recognition module 112 recognizes the user's voice, enabling voice interaction between the mobile terminal and the user. The system settings module 113 provides a settings interface that allows the user to configure personalized configuration information for abnormality monitoring.
[0044] The monitoring and analysis system 120 may include: a baseline information generation module 121, an abnormal image change identification module 122, a flame identification module 123, a face recognition module 124, a stranger identification module 125, a fall identification module 126, and an animal identification module 127. The baseline information generation module 121 is used to generate baseline information based on the surveillance video image. The abnormal image change identification module 122 is used to monitor whether the image has abnormal changes. The flame identification module 123 is used to identify whether there is a flame in the surveillance video image and, if so, issues an early warning message to alert the user of a fire. The face recognition module 124 is used to identify whether the person in the surveillance video image is the user or a user-defined authorized user (e.g., someone living with the user). If the person is identified as the user or a authorized user, it notifies the fall identification module 126 to detect whether the user or the authorized user has fallen. If the person is not identified as the user or a authorized user, that is, if the person is identified as a stranger, it notifies the stranger identification module 126 to issue an early warning message to alert the user of the intrusion. The stranger identification module 125 is used to issue an early warning message to warn the user of the intrusion of a stranger when a person in the surveillance video image is identified as a stranger. The fall detection module 126 is used to detect whether the user or the legitimate user has fallen when the person is identified as the user himself or the legitimate user; if the user himself or the legitimate user has fallen, an early warning message is issued to warn the user himself or the legitimate user of the fall, and / or to notify relevant personnel or institutions (such as the user's family, hotel staff, police, hospitals) of the user's fall. The animal identification module 127 is used to identify whether there are animals in the surveillance video image; if an animal is identified, an early warning message is issued to warn the user of the intrusion of an animal.
[0045] It should be understood that in this application Figure 2 The system shown is only an example. In actual applications, in addition to the flame recognition module 123, face recognition module 124, stranger recognition module 125, and animal recognition module 127 mentioned above in this application, other abnormality recognition modules can also be deployed according to actual conditions, such as a flood recognition module, a mudslide recognition module, etc.
[0046] The hardware system 130 may include: a first camera 131, a second camera 132, a network module 133, a memory 134 and a power supply 135. Among them, the first camera 131 and the second camera 132 are used to capture video images. The network module 133 is used to provide a wired or wireless network so that the mobile terminal 100 can communicate with the device 200 through the network module 133. The memory 134 is used to store monitoring video images and information related to abnormal monitoring, such as abnormal events and videos related to abnormal events, interactive information, and early warning information. The power supply 135 is used to provide power for the mobile terminal 100. The power supply 135 has a charging interface, which can be connected to other external power sources through the charging interface to charge the mobile terminal 100. When the mobile terminal 100 remains connected to the external power source through the charging interface, the mobile terminal 100 can continue to execute the abnormal monitoring method.
[0047] In the embodiment of the present application, the first camera 131 has a larger field of view (VOF) than the second camera 132, so that the coverage of the image captured by the first camera 131 is larger than the coverage of the image captured by the second camera 132. The pixels of the image captured by the first camera 131 are smaller than the pixels of the image captured by the second camera 132, so that the image captured by the second camera 132 is of higher quality and clearer than the image captured by the first camera 131.
[0048] In some embodiments, the first camera 131 may be a wide-angle camera, used to expand the scanning range of the monitoring area and provide feedback and warning of abnormalities within the monitoring area. The second camera 132 may be the main camera of the mobile terminal 100, used to capture images of abnormalities, identify the abnormal area, and respond accordingly.
[0049] See also Figure 3 and Figure 4 The mobile terminal 100 includes a first surface 140 and a second surface 150. The first surface 140 refers to the surface where the screen or internal screen of the mobile terminal is located, and the second surface 150 refers to the surface of the mobile terminal without a screen or where the external screen is located. The first camera 131 and the second camera 132 can be set on the first surface 140 of the mobile terminal 100.
[0050] See also Figure 3 Taking a foldable phone as an example, the first surface refers to the surface of the phone without a screen, or the surface where the outer screen of the phone is located. In contrast, the surface with a screen, or the surface where the inner screen of the phone is located, is called the second surface.
[0051] See also Figure 4Taking the mobile terminal being a straight-screen mobile phone as an example, the first surface may refer to the back of the straight-screen mobile phone (the surface away from the user or the surface facing away from the user when the user uses the straight-screen mobile phone). Conversely, the second surface may refer to the back of the straight-screen mobile phone (the surface facing the user when the user uses the straight-screen mobile phone).
[0052] In some embodiments, the mobile terminal 100 may further include other cameras, for example, a telephoto camera mounted on the first surface of the mobile terminal 100 and a selfie camera mounted on the second surface of the mobile terminal 100, so as to provide the user with a good photo-taking experience.
[0053] Device 200 may be an electronic device or a server. The server may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0054] It should be noted that the abnormality monitoring method of the present application is executed when the mobile terminal is in the monitoring position. Figure 5 , when the mobile terminal 100 is in the monitoring position, the first camera and the second camera are not blocked and can capture the doors and windows of the room where the mobile terminal is located. The fan-shaped shadow part is the monitoring scanning range when the mobile terminal 100 is in the monitoring position. It is understandable that the monitoring position can be set by the user according to the actual environment. The user can use a mobile phone holder or a gimbal or similar item to place the mobile terminal folding screen mobile phone in the monitoring position. In particular, if the mobile terminal is a folding screen mobile phone, such as Figure 6 and Figure 7 As shown, the user can open the foldable screen mobile phone so that the foldable screen mobile phone is in hover mode, that is, the foldable screen is in a semi-folded state, and then place the foldable screen mobile phone in the monitoring position, so as to use the foldable screen mobile phone for abnormal monitoring.
[0055] See also Figure 8 , Figure 8 FIG. 1 is a flow chart of an abnormality monitoring method provided in an embodiment of the present application. The abnormality monitoring method may include steps S110 to S140.
[0056] Step S110: Acquire a first video image captured by the first camera and a second video image captured by the second camera.
[0057] The user can activate the monitoring function by touching a touch button on the abnormality monitoring interface of the mobile terminal or by voice. In response to the activation of the monitoring function, the mobile terminal calls the first camera and the second camera, and the first camera starts to capture the first video image and the second camera starts to capture the second video image.
[0058] As mentioned above, the viewing angle of the first camera is larger than that of the second video image, so that the coverage area of the first video image is larger than that of the second video image. The pixel of the second camera is lower than that of the second camera, so that the second video image is clearer than the first video image.
[0059] When the mobile terminal is at the monitoring position, a first video image captured by the first camera and a second video image captured by the second camera are acquired.
[0060] In some embodiments, the user can manually determine whether the mobile terminal is in a monitoring location. For example, the user can place the mobile terminal in a user-defined monitoring location and then activate the first and second cameras of the mobile terminal to capture images or record videos. If the first and second cameras can capture the doors and windows of the room where the mobile terminal is located, the mobile terminal can be determined to be in the monitoring location. The user can then activate the abnormal monitoring function. In response to the activation of the abnormal monitoring function, the user can obtain a first video image captured by the first camera and a second video image captured by the second camera.
[0061] In some embodiments, the mobile terminal may determine whether it is in a monitoring location (see steps S210 to S220 below for details). In response to determining that the mobile terminal is in a monitoring location, the mobile terminal may automatically trigger an abnormality monitoring function, turning on the first camera and the second camera to obtain a first video image captured by the first camera and a second video image captured by the second camera.
[0062] In some embodiments, considering that images captured in low-light environments such as at night or on rainy days are relatively dark and inconvenient for image processing, after acquiring the first and second video images, dark-light processing can be performed on the first and second video images to enhance visibility and contrast of the first and second video images, while repairing noise, artifacts, color distortion, etc. hidden in darkness or introduced by increasing brightness. This allows subsequent steps 120 to S130 to use the dark-light-processed first and second video images for abnormality monitoring and analysis, reducing the difficulty of subsequent abnormality monitoring and analysis and improving the accuracy of subsequent abnormality monitoring and analysis. The dark-light processing method can refer to related technologies, such as histogram equalization methods, gamma correction methods, methods based on retinal theory that use the image's reflectance map as a dark-light-enhanced image, and deep learning methods based on convolutional neural networks.
[0063] Step S120: detecting whether abnormal changes occur in the monitoring image based on the multiple frames of the first video image.
[0064] The multiple frames of the first video image include: the current frame (detection frame) in the first video image and multiple frames (e.g., 10 frames) preceding the current frame. The number of multiple frames can be set in advance based on the accuracy requirements of the monitoring image anomaly analysis.
[0065] The monitoring screen refers to the first video image, that is, the image captured by the first camera. Abnormal changes in the monitoring screen may include: the appearance of new objects (e.g., people, animals, flames, etc.) in the current frame image and / or a significant change in the objects in the current frame image compared to the previous frame image (e.g., a person in the image falls). It is understood that when the current frame image is static or only has normal light and shadow changes, it can be considered that the monitoring screen has not changed, that is, the monitoring screen has not undergone abnormal changes.
[0066] Specifically, reference information may be generated based on multiple frames of images preceding the current frame of the first video image; and whether abnormal changes occur in the monitoring screen may be detected based on the current frame of the image and the reference information.
[0067] Among them, multiple frames of images before the current frame image are used to generate reference information corresponding to the current frame image, and the reference information is used to determine whether abnormal changes occur in the current frame image, thereby achieving preliminary abnormality detection.
[0068] In the embodiment of the present application, the reference information may include a reference image and / or a reference feature vector. The reference image may be calculated as an average image of multiple frames preceding the current frame image, and / or the reference feature vector may be calculated as an average value of feature vectors of multiple frames preceding the current frame image.
[0069] Taking the 10 frames before the current frame as an example, the method of generating the reference image of the current frame is as follows: obtain the 10 frames before the current frame; calculate the average image of the 10 frames; and use the average image as the reference image of the current frame.
[0070] Taking the 10 frames preceding the current frame as an example, the baseline feature vector for the current frame is generated as follows: a feature extractor (encoder) is used to extract feature vectors for each of the 10 frames preceding the current frame; the average of these 10 frames' feature vectors is calculated; and this average is used as the baseline feature vector for the current frame. The feature extractor employs a convolutional neural network (CNN), such as the Residual Neural Network (ResNet) or the deep convolutional neural network AlexNet.
[0071] An image difference method and / or an image feature distance detection method may be used to detect whether abnormal changes occur in the monitoring image based on the current frame image and the reference information.
[0072] In some embodiments, using an image difference method, based on the current frame image and reference information, detecting whether an abnormal change has occurred in the surveillance image may include: performing a difference operation on the reference image and the current frame image to obtain a difference value; if the difference value is greater than a difference threshold, determining that an abnormal change has occurred in the surveillance image; if the difference value is less than or equal to the difference threshold, determining that no abnormal change has occurred in the surveillance image. The difference threshold can be set in advance based on the actual accuracy requirements for image anomaly analysis; for example, the difference threshold can be set to 0.8.
[0073] In other embodiments, an image feature distance detection method is used. Based on the current frame image and the reference information, detecting whether an abnormal change has occurred in the monitoring image may include: extracting the feature vector of the current frame image; calculating the distance (e.g., cosine distance or Euclidean distance) between the feature vector of the current frame image and the reference feature vector; if the distance is greater than a distance threshold, determining that an abnormal change has occurred in the monitoring image; if the distance is less than or equal to the distance threshold, determining that no abnormal change has occurred in the monitoring image. The distance threshold can be set in advance based on the actual accuracy requirements for image anomaly analysis.
[0074] In some other embodiments, the image feature distance detection method and the image feature distance detection method are used to detect whether abnormal changes have occurred in the monitoring screen based on the current frame image and the reference information, which may include: performing a differential operation on the reference image and the current frame image to obtain a differential value, and detecting whether the differential value is greater than a differential threshold; extracting the feature vector of the current frame image; calculating the distance between the feature vector of the current frame image and the reference feature vector, and detecting whether the distance is greater than a distance threshold; if the differential value is greater than the differential threshold or the distance is greater than the distance threshold, determining that abnormal changes have occurred in the monitoring screen; if the differential value is less than or equal to the differential threshold and the distance is less than or equal to the distance threshold, determining that no abnormal changes have occurred in the monitoring screen.
[0075] If an abnormal change in the surveillance image is detected based on the current frame image, the process proceeds to step S130, where further abnormality analysis is performed based on the second video image. If no abnormal change in the surveillance image is detected based on the current frame image, the process proceeds to the next frame image to detect abnormal changes in the surveillance image, in the same manner as the current frame image detection method, until an abnormal change in the surveillance image is detected or the abnormality monitoring function is disabled.
[0076] Step S130: When an abnormal change occurs in the monitoring image, an abnormal event is detected based on the second video image.
[0077] When an abnormal change occurs in the surveillance footage, the image type of the second video image is identified; based on the image type, an abnormal event is determined. Image types may include images containing people, images containing flames, and images containing animals. It is understood that in actual applications, the image type can be set based on actual needs. For example, in addition to the aforementioned image types, image types may also include images containing water, images containing mud and sand, and so on.
[0078] Specifically, the second video image can be input into a pre-trained abnormal image classification model to obtain the image type output by the abnormal image classification model. The abnormal image classification model is pre-trained using a convolutional neural network with a depth of 50 layers (referred to as a ResNet-50 network), whose input is the second video image and output is the image type corresponding to the second video image. For example, a training data set including video image samples of unlabeled image types and video image samples of labeled image types can be obtained; the unlabeled image type is used as input and the labeled image type video image samples are used as output to train the ResNet-50 network multiple times; the loss value of each training is calculated, and the training is stopped when the loss value is less than the loss threshold; the trained ResNet-50 network is deployed as an abnormal image classification model in a mobile terminal.
[0079] In the embodiment of the present application, determining an abnormal event based on the image type may include:
[0080] ① If the image type is an image containing flames, the abnormal event is determined to be a fire.
[0081] ② If the image type is an image containing animals, the abnormal event is determined to be an animal intrusion.
[0082] ③ If the image type is an image containing a person and the person in the second video image is a stranger, the abnormal event is determined to be a stranger breaking in.
[0083] ④ If the image type is an image containing a person and the person in the second video image falls, it is determined that the abnormal event is the user falling.
[0084] It is understandable that the above-mentioned abnormal image classification model can output at least one image type, and the above-mentioned situations ① to ④ may only occur in one or more cases at the same time. That is, in an embodiment of the present application, abnormal events may include: one or more combinations of fire, animal intrusion, stranger intrusion, and user fall. Among them, the priority of the fourth case is lower than the priority of the other three abnormal events. In some embodiments, the priorities of ① to ④ can be set in descending order, so that when multiple abnormal events occur, the more dangerous abnormal events are given priority warning.
[0085] In an embodiment of the present application, when the image type is an image containing a person, face detection can be performed on the second video image; and abnormal events of the above categories ③ and / or ④ are determined based on the face detection results, thereby improving the accuracy of abnormal warning.
[0086] Specifically, when the second video image is an image including a person, and there is a face in the second video image, a face detection algorithm can be used to extract the face in the second video image; and through face recognition, the user identity of the person in the second video image can be determined. Specifically, the identified user identity can be compared with the user identity pre-stored in the mobile terminal (for example, the identified face can be compared with the face pre-stored in the mobile terminal) to determine whether there is a stranger in the person in the second video image; if the user identities pre-stored in the mobile terminal include all the identified user identities, the face detection result is determined to be that there is no stranger in the person in the second video image; if the user identities pre-stored in the mobile terminal do not completely include all the identified user identities, the face detection result is determined to be that there is a stranger in the person in the second video image. In some embodiments, when a stranger is identified, the stranger can be marked in the second video image, and the image with the stranger mark can be stored and displayed, so as to facilitate viewing of the stranger and its related activity tracks.
[0087] The user identities pre-stored in the mobile terminal may include the user and other legally added users. Other legally added users may include, but are not limited to, people living with the user or the user's family, friends, and close friends. Optionally, the user identities pre-stored in the mobile terminal may be stored as specific user identities, such as "myself," "friends," or "family." Alternatively, the user identities may be stored in the form of faces, i.e., faces of legally added users, such as the user's own face, faces of friends, or faces of family.
[0088] In some embodiments, if the face detection result indicates that a stranger is among the people in the second video image, the abnormal event is determined to be an intrusion by a stranger.
[0089] In some embodiments, if the face detection result indicates that there are no strangers among the people in the second video image, it is determined that there is no abnormal event of a stranger breaking in.
[0090] In some embodiments, taking into account the personal danger posed by special users (for example, elderly users or users with special physical conditions who cannot fall) after falling, when the face detection result indicates that there are no strangers among the characters in the second video image, the human joints in the second video image can be further extracted; based on the human joints, it is determined whether the user has fallen (user falling is an abnormal event), so that a personal safety warning can be issued in a timely manner after it is determined that the special user has fallen.
[0091] The human joint points are the points marked to represent the positions of human joints. For example, Figure 9 As shown, the human body joints may include: head joints 14, 15, 16, 17, neck joint 0, shoulder joints 2, 1, 5, elbow joints 3, 6, hand joints 4, 7, waist joints 8, 11, leg joints 9, 12, and foot joints 10, 13.
[0092] This application is based on anomaly detection on mobile terminals. Taking into account the limited computing resources, in an embodiment of this application, a human posture estimation algorithm based on deep learning is adopted to evaluate the posture of the person in the second video image and obtain the human joint points (and their coordinates).
[0093] The human pose estimation algorithm based on deep learning can be an OpenPose human pose estimation network. The second video image is input into the OpenPose human pose estimation network, and the OpenPose human pose estimation network can output the human joint points and their coordinates of the characters in the second video image. As an example, Figure 9 It is a schematic diagram of the human joint points of a person (such as the user himself) output by the OpenPose human pose estimation network. Figure 9 The characters in the game include 18 points of articulation from 0 to 17.
[0094] Specifically, the OpenPose human pose estimation network first detects the limb joints of all characters in the second video image, and then assembles the joints of the same character. The overall network structure of the OpenPose human pose estimation network uses the Visual Geometry Group (VGG) network as a pre-trained framework, where the VGG network is used to extract convolutional features from the image. The OpenPose human pose estimation network is divided into two parts, both of which can simultaneously predict confidence maps for the extracted key points, encode the correlation vector field between adjacent key points, and regress the confidence map and correlation vector field respectively.
[0095] like Figure 10 The OpenPose human pose estimation network architecture shown in Figure 1 is used. F is the output of the VGG network. Branch 1 predicts the confidence map S (representing the prediction of the joints themselves), and branch 2 predicts the association vector field L (representing the association between adjacent key points). The 3×3, 1×1, and 7×7 kernel sizes in the convolutional layers represent the corresponding convolutional kernel sizes. Each regression of S and L completes one round of iterative prediction. After t (t ≥ 2) consecutive rounds of iteration, the entire prediction network architecture is formed. At each stage, the losses f1 and f2 are calculated, and S, L, and F (the original input) are concatenated to obtain the input for the next stage of prediction training. After multiple rounds of iteration, S can play a certain role in distinguishing the left and right structures of the prediction network system. The degree of differentiation increases with the number of iterations.
[0096] Determining abnormal events based on human body joints may include: determining the moving speed of the human body's center of mass based on the human body joints; detecting whether the user falls based on the moving speed to obtain a first detection result; detecting whether the user falls based on the height of the neck joint to obtain a second detection result; detecting whether the user falls based on the relative position relationship between the shoulder joint and the waist joint to obtain a third detection result; if the first detection result, the second detection result, and the third detection result all indicate that the user has fallen, determining that the abnormal event is a user fall; if one of the first detection result, the second detection result, and the third detection result indicates that the user has not fallen, determining that there is no abnormal event of the user falling.
[0097] Detecting whether the user falls based on the moving speed, obtaining the first detection result may include: calculating the average value of the human body joints and using the average value as the center of mass of the human body; calculating the moving speed of the center of mass of the human body; if the moving speed is greater than or equal to the speed threshold, determining that the first detection result is that the user has fallen; if the moving speed is less than the speed threshold, determining that the first detection result is that the user has not fallen.
[0098] In some embodiments, in order to ensure the stability of anomaly detection, the neck joint points (such as Figure 9 The joint point 0 in the middle of the shoulder (such as Figure 9 The joint point 1) and the joint point on one side of the waist (such as Figure 9 The average value of the coordinates of the joint points 8) in the figure is used as the coordinates of the center of mass of the human body.
[0099] In some embodiments, the moving speed of the human body center of mass can be calculated once every second video image with a specified number of frames. The specified number of frames can be set according to actual needs. For example, the specified number of frames can be 10 frames. For example, expression (1) can be used to calculate the moving speed of the human body center of mass once every m frames of the second video image, based on the neck joint point (such as Figure 9 The joint point 0 in the middle of the shoulder (such as Figure 9 The joint point 1) and the joint point on one side of the waist (such as Figure 9 Calculate the moving speed of the human body's center of mass at the joint point 8 in the figure:
[0100]
[0101] Among them, v represents the moving speed of the center of mass of the human body; represents the ordinate of the i-th joint point in the second video image of the m-th frame; represents the vertical coordinate of the i-th joint point in the second video image of the first frame; t m ―t1 represents the time from the 1st frame of the second video image to the mth frame of the second video image.
[0102] In some embodiments, the speed threshold can be set in advance based on multiple experimental measurement data. For example, the speed threshold can be calculated using expression (2) based on the data of M experimenters simulating N times of slow human falls:
[0103]
[0104] Among them, v th represents the speed threshold; i represents the i-th person, and j represents the j-th experiment.
[0105] As an example, according to expression (2), assuming that the speed threshold is set based on the data of 5 experimenters simulating 15 slow falls of the human body, with M = 5 and N = 15, the speed threshold v can be set as th Set to 125 pixels / second (pixel / s), 125pixel / s means that the user falls 125 pixels every 1 second.
[0106] In some embodiments, detecting whether a user has fallen is performed based on the height of the neck joint point, and obtaining a second detection result may include: obtaining the height (vertical coordinate) of the neck joint point of a (legal) person in the second video image; if the height of the neck joint point exceeds a preset fluctuation range, determining that the second detection result is that the user has fallen; if the height of the neck joint point is within the preset fluctuation range, determining that the second detection result is that the user has not fallen.
[0107] The pre-set fluctuation range may be the fluctuation range of the height (ordinate) of the neck joint when the user is in normal activity. For example, the pre-set fluctuation range may include: the fluctuation range of the height (ordinate) of the neck joint when the human body is standing normally, the fluctuation range of the height (ordinate) of the neck joint when the human body is sitting normally, the fluctuation range of the height (ordinate) of the neck joint when the human body is lying flat on the bed, etc.
[0108] In other embodiments, detecting whether the user has fallen based on the height of the neck joint may include: calculating the height difference between the neck joint and the ground; if the height difference between the neck joint and the ground is less than or equal to a height difference threshold, determining that the second detection result is that the user has fallen; and if the height difference between the neck joint and the ground is greater than the height difference threshold, determining that the second detection result is that the user has not fallen. The height difference threshold may be set in advance based on actual needs or experimental data.
[0109] In some embodiments, detecting whether the user falls is based on the relative position relationship between the shoulder joint and the waist joint, and obtaining the third detection result may include: detecting whether the height of the shoulder joint is equal to or lower than the height of the waist joint; if the height of the shoulder joint is equal to or lower than the height of the waist joint, determining that the third detection result is that the user has fallen; if the height of the shoulder joint is higher than the height of the waist joint, determining that the third detection result is that the user has not fallen.
[0110] In an embodiment of the present application, based on the preliminary detection of abnormal changes in the monitoring screen based on the first video image, the abnormal event is further determined based on the second video image, which can improve the accuracy of determining the abnormal event and facilitate subsequent accurate warning of the abnormal event.
[0111] Step S140: issuing a warning message based on the abnormal event.
[0112] As mentioned above, abnormal events may include: a fire, an animal breaking in, a stranger breaking in, or a user falling down, or one or more combinations thereof.
[0113] If the abnormal event includes a fire, a warning message is issued to warn of the fire. If the abnormal event includes an animal intrusion, a warning message is issued to warn of the animal intrusion. If the abnormal event includes a stranger intrusion, a warning message is issued to warn of the stranger intrusion. If the abnormal event includes a user falling, a warning message is issued to warn of the user falling.
[0114] Among them, the methods of issuing warning information may include but are not limited to photoelectric alarm prompts, calling relatives, sending text messages to relatives, directly calling the alarm number and sending the user's current location (current location of the mobile terminal), etc. The specific method of issuing warning information can be customized in advance by the user based on the user interaction system of the mobile terminal, and the embodiments of this application do not limit the specific method of issuing warning information.
[0115] Based on steps S110 to S140, the first camera of the mobile terminal is used to capture the first video image, and the second camera is used to capture the second video image. Due to the wider viewing angle of the first camera, the use of multiple frames of the first video image can more quickly detect whether there has been an abnormal change in the monitoring image. When an abnormal change in the monitoring image is detected, the second video image is used to detect the abnormal event, which can accurately determine the abnormal event and achieve precise early warning of the abnormal event, ensuring the personal safety of the user. This expands the application of mobile terminals, making it possible to implement abnormal monitoring and early warning through mobile terminals, and can protect the personal safety of users at a low cost.
[0116] See also Figure 11 , Figure 11 This is a partial flow chart of an abnormality monitoring method provided by another embodiment of the present application. After the monitoring function is enabled and before step S120, the abnormality monitoring method may include steps S210 to S240.
[0117] Step S210: Detect whether the mobile terminal is in a stationary state.
[0118] The mobile terminal being in a stationary state may refer to: ① the mobile terminal being in a stationary state; ② the mobile terminal being in a stationary state for a period of time.
[0119] In some embodiments, the mobile terminal is a folding screen mobile terminal, and it is possible to first detect whether the folding screen of the mobile terminal is in a semi-folded state; when the folding screen of the mobile terminal is in the semi-folded state (such as Figure 6 and Figure 7 ), detect whether the mobile terminal is in a stationary state; when the folding screen of the mobile terminal is not in a semi-folded state (as shown Figure 4 As shown), continue to detect whether the folding screen of the mobile terminal is in a semi-folded state and repeat the above detection operation until it is detected that the folding screen of the mobile terminal is in a semi-folded state.
[0120] In some embodiments, an inertial measurement unit (IMU) in the mobile terminal can be used to detect whether the mobile terminal is stationary. For example, data output by the IMU can be obtained and the acceleration of the mobile terminal can be calculated based on the data. When the acceleration of the mobile terminal is zero, the mobile terminal is determined to be stationary. When the acceleration of the mobile terminal is not zero, data output by the IMU is continuously obtained and the above detection operation is repeated until the mobile terminal is detected to be stationary.
[0121] In other embodiments, the first camera (or second camera) of the mobile terminal is called to obtain multiple frames of images captured by the first camera (or second camera); it is detected whether the pictures of the multiple frames of images are the same; when the pictures of the multiple frames of images are the same, it is determined that the mobile terminal is in a stationary state; when the pictures of the multiple frames of images are different, the multiple frames of images captured by the first camera (or second camera) continue to be obtained and the above detection operation is repeated until it is detected that the mobile terminal is in a stationary state.
[0122] In some other embodiments, a prompt message (such as Figure 12 In response to the monitoring function being turned on, the mobile terminal is in a stationary state by default; or in response to negative feedback information input by the user, it is determined that the mobile terminal is not in a stationary state.
[0123] Among them, the prompt information can be Figure 1 The monitoring screen display module shown is output and displayed on the screen of the mobile terminal; and / or the prompt information can also be displayed through Figure 1 The speech recognition module shown outputs in audio form.
[0124] Among them, the monitoring function is enabled by the user, such as Figure 12 As shown, users can turn on the monitoring function by touching the monitoring function button on the mobile terminal, or by speaking Figure 1 The voice recognition module shown interacts to enable monitoring functionality.
[0125] Among them, negative feedback information indicates that the monitoring function is not turned on, such as Figure 12 As shown, the user can input negative feedback information by touching the cancel button on the mobile terminal, or by speaking Figure 1 The speech recognition module shown interacts to input negative feedback information.
[0126] It can be understood that step S210 can be executed throughout the entire process of the abnormal monitoring method to determine whether the mobile terminal moves during the abnormal monitoring process. If it is detected that the mobile terminal moves, the operations of steps S210 to S240 and steps S120 to S140 need to be repeated. For example, if the mobile terminal moves during the abnormal monitoring process, the baseline information in the above step S120 needs to be regenerated.
[0127] Step S220: When the mobile terminal is in a stationary state, detecting whether the mobile terminal is in a monitoring location.
[0128] like Figure 5 As shown, when the mobile terminal is in the monitoring position, the shooting range of the first camera can cover the doors and windows of the room where the mobile terminal is located. The monitoring position can be arranged by the user according to the specific environment of the room where the mobile terminal is located. In some embodiments, in order to facilitate the user to place the mobile terminal in the monitoring position, before step S210, the mobile terminal can first sense the layout of the room where the mobile terminal is located through various sensors thereon, and obtain a layout diagram of the room where the mobile terminal is located; generate an optimal monitoring position based on the layout diagram; display the layout diagram and the optimal monitoring position (such as Figure 13 ), prompting the user to place the mobile terminal in the best monitoring position, thereby simplifying the user's operation difficulty and reducing the time of placing the mobile terminal. After the user places the mobile terminal in the best monitoring position, the user can touch the completed button on the mobile terminal (as shown Figure 13 As shown) or by voice input confirmation information, in response to the confirmation information, step S210 is started.
[0129] When the mobile terminal is in a stationary state, detect whether the first camera and the second camera of the mobile terminal are blocked; when the first camera and the second camera are not blocked, detect whether the shooting range of the first camera of the mobile terminal covers the door and window range of the room where the mobile terminal is located; when the shooting range of the first camera covers the door and window range, determine that the mobile terminal is in a monitoring position.
[0130] When one of the first camera and the second camera is blocked, a prompt message (such as Figure 14 After receiving the prompt, the user can remove the camera obstruction by moving the mobile terminal or removing the obstructing object on the camera, and then touch the resolved button on the mobile terminal (as shown) to clear the camera obstruction. Figure 14 In response to the confirmation message, the user can continue to detect whether the first camera and the second camera of the mobile terminal are blocked, or directly determine that the mobile terminal is in the monitoring position.
[0131] When the shooting range of the first camera does not cover the door and window range, a prompt message (such as Figure 15 ) or through voice prompts, prompting that the current shooting range does not cover the door and window range. After receiving the prompt, the user can adjust the camera by moving the mobile terminal or adjusting the angle of the mobile terminal, and then touch the adjusted button on the mobile terminal (such as Figure 15 In response to the confirmation message, the mobile terminal can continue to detect whether the shooting range of the first camera of the mobile terminal covers the door and window range of the room where the mobile terminal is located, until it is detected that the shooting range of the first camera covers the door and window range (as shown in the figure). Figure 5 As shown), it is determined that the mobile terminal is in the monitoring position.
[0132] Step S230 (see step S110 ): when the mobile terminal is at the monitoring position, a first video image captured by the first camera and a second video image captured by the second camera are acquired.
[0133] Step S240: Displaying the first video image and the second video image, wherein the second video image is displayed on the first video image, and the second video image does not completely cover the first video image.
[0134] See also Figure 16 and Figure 17 , the second video image can be displayed floating on the first video image.
[0135] In some embodiments, such as Figure 16 As shown, the mobile terminal is a foldable screen mobile phone and the foldable screen of the foldable screen mobile phone is in a semi-folded state (the foldable screen mobile phone is in hovering mode at this time), and the virtual foldable screen crease line divides the foldable screen into two parts of the screen, one part of the screen (the screen above the virtual foldable screen crease line with a shadow) is perpendicular to the ground, and the other part of the screen (the screen below the virtual foldable screen crease line) is parallel to the ground. The first video image can be displayed on the part of the screen perpendicular to the ground, and the second video image is displayed in suspension on the first video image. In addition, a suspension setting button can be displayed on the part of the screen parallel to the ground. When the user clicks the suspension setting button, the setting interface will be entered. The user can make relevant settings for the abnormal monitoring system based on the setting interface, for example, set the cropping frame rate of the camera, the corresponding response operations when different abnormal situations occur, the entry and exit of the abnormal monitoring system, the method of turning on the monitoring function, and other operations.
[0136] In some embodiments, such as Figure 17As shown, if the mobile terminal is a foldable screen phone and the inner screen of the foldable screen phone is in a fully unfolded state, the first video image can be displayed on the entire inner screen, and the second video image and the aforementioned floating setting button can be displayed floatingly on the first video image, without overlapping the floating setting button and the second video image. If the mobile terminal is not a foldable screen phone, the first video image can be displayed on the entire screen of the mobile terminal, and the second video image and the aforementioned floating setting button can be displayed floatingly on the first video image, without overlapping the floating setting button and the second video image.
[0137] In some embodiments, a user can touch the second video image. In response to the user's touch operation on the second video image, the mobile terminal may not display the first video image, but instead magnify and display the second video image at the location where the first video image was originally displayed, while keeping the floating settings button stationary, thereby facilitating the user's viewing of the second video image. After the second video image is magnified and displayed, the user can touch an object (e.g., a person) in the second video image. In response to the touch operation on the object in the second video image, the mobile terminal may magnify and display the object, facilitating the user's viewing of the object. The user can then perform a designated return operation on the screen, causing the mobile terminal to control the screen to return to the previous screen in response to the designated return operation. For example, after magnifying an object, the mobile terminal may magnify and display the second video image in response to the designated return operation. For another example, after magnifying and displaying the second video image, the mobile terminal may simultaneously display the first and second video images in response to the designated return operation. The designated return operation may be customized by the user in advance through a settings interface or may be a system default operation, such as swiping from the right side of the screen to the left side.
[0138] Based on steps S210 to S240, an optimal monitoring position can be automatically generated, and the user can be guided to place the mobile terminal in the optimal monitoring position, thereby simplifying the user's operation and reducing the time required to place the mobile terminal. When the mobile terminal is in the monitoring position, the monitoring screen is acquired and displayed, allowing the user to view the monitoring screen and monitoring screen details in real time. In addition, a floating setting button and setting interface are provided, allowing the user to customize the relevant settings for the abnormal monitoring system, which can enhance the user experience.
[0139] See also Figure 18 , Figure 18 Another embodiment of the present application provides a flowchart of an abnormality monitoring method. After the monitoring function is enabled, the abnormality monitoring method may include steps S310 to S3200.
[0140] Step S310: Detect whether the mobile terminal is moving.
[0141] If the mobile terminal is moving, it is determined that the mobile terminal is not stationary. The captured environmental image during movement is changing, making it difficult to generate a base environmental image (i.e., a reference image). This is because the base environmental image, which serves as the standard for determining whether the environmental image (i.e., the monitoring image) has undergone abnormal changes, should remain essentially unchanged. Therefore, while the mobile terminal is moving, step S310 can be continued to detect whether the mobile terminal has moved.
[0142] If the mobile terminal does not move, it is determined that the mobile terminal is in a stationary state and the environment picture can be collected. In order to collect the complete environment picture, step S320 can be entered first to detect whether the camera of the mobile terminal is blocked.
[0143] It is understandable that, in order to ensure the stability of subsequent environmental image acquisition and the accuracy of the environmental basic image, step S310 is kept running from the time the monitoring function is turned on until the monitoring function is turned off.
[0144] Step S320: Detect whether the camera is blocked.
[0145] When the mobile terminal is in a stationary state, it is detected whether the first camera and the second camera of the mobile terminal are blocked, so as to ensure that the complete first video image and the second video image are captured when both the first camera and the second camera are not blocked.
[0146] If the first camera and / or the second camera is blocked, the process proceeds to step S330 to prompt the user that the camera is blocked, so that the user can move the mobile terminal or remove the object blocking the camera.
[0147] If both the first camera and the second camera are not blocked, the process proceeds to step S340 to capture the environment image.
[0148] It is understandable that, in order to ensure the integrity of subsequent environmental image acquisition, step S320 is kept running from the beginning until the monitoring function is turned off.
[0149] Step S330: Prompt that the camera is blocked.
[0150] Considering that the user may remove the camera from being blocked by moving the mobile terminal or removing the object blocking the camera, after prompting the user that the camera is blocked, the process can return to step S310 to continue detecting whether the mobile terminal has moved.
[0151] Step S340: Collect environmental images.
[0152] When the mobile terminal is not moving and the camera is not blocked, the environment picture is collected. Specifically, the first camera can be used to collect the first video image and the second camera can be used to collect the second video image. The picture of the first video image and the second video image is the environment picture.
[0153] Step S350: Generate an environment basic screen.
[0154] After capturing multiple frames of the first video image, a reference image can be generated based on these multiple frames. The reference image serves as the baseline image for the environment and can subsequently serve as a criterion for determining whether any abnormal changes have occurred in the image. It is understood that the reference image is dynamically updated. That is, for each current frame, the reference image corresponding to the current frame is generated using multiple frames of video imagery preceding the current frame. Because the mobile terminal remains unchanged, the reference image obtained after the dynamic update of the reference image is virtually identical to the previous reference image, with minimal differences. Therefore, the baseline image does not significantly change.
[0155] Step S360: Detect whether the environment basic screen has changed.
[0156] It is understandable that if the mobile terminal is a foldable screen mobile terminal, the foldable screen mobile terminal is basically placed flat on the desktop without the support of the gimbal, so the camera cannot be automatically rotated, so the basic environment picture is basically fixed. The basic environment picture will only change when the user moves the mobile terminal, and the basic environment picture will be triggered to readjust. Based on this, step S360 is designed to determine whether the basic environment picture has changed, so that the basic environment picture can be dynamically updated in time after the mobile terminal moves. It is understandable that whether the basic environment picture has changed here refers to the change triggered by the movement of the mobile terminal, which refers to a larger change compared to the previous basic environment picture, rather than the change in the basic environment picture caused by the dynamic update of the reference image as described above.
[0157] If the environment basic screen is changed, the process returns to step S350 to regenerate the environment basic screen.
[0158] If the basic environment image has not changed, the process proceeds to step S360 , scanning the environment for 1 second per frame, and acquiring the environment image once every frame.
[0159] Step S370: Scan the environment for 1 second per frame.
[0160] The first video image captured by the first camera and the second video image captured by the second camera may be acquired once every frame.
[0161] Step S380: Detect whether it is a dark environment.
[0162] Whether the current environment is a dark environment can be detected based on the first video image and / or the second video image. It is understandable that since the first video image covers a wider range and has lower pixels than the second video image, the first video image can be used for dark environment detection, thereby improving detection efficiency.
[0163] If the current environment is a dark environment and the captured image is dark, in order to improve the accuracy of abnormal monitoring and reduce the difficulty of calculation, step S390 can be entered to perform dark light enhancement processing on the first video image and the second video image to improve the brightness of the captured image.
[0164] If the current environment is not a dark environment, the brightness of the collected image has little impact on the accuracy of abnormal monitoring and the difficulty of calculation. At this time, step S3100 can be entered to compare the current picture with the basic environment picture.
[0165] Step S390: Image dark light enhancement processing.
[0166] The first video image and the second video image may be subjected to dark light enhancement processing to increase the brightness of the captured image, thereby improving the accuracy of abnormal monitoring and reducing the difficulty of calculation. After executing step S390, step S3100 may be entered to compare the current image with the environmental base image.
[0167] Step S3100: Compare the current image with the environment basic image.
[0168] The current screen may be the current first video image screen and / or the second video image screen. It is understandable that, since the first video image has a wider coverage area and lower pixel count than the second video image, the first video image screen can be compared with the environment base screen to improve detection efficiency. Specifically, the difference between the current screen and the environment base screen (e.g., the difference value or distance between the two) can be calculated, and then step S3110 is entered to detect whether the monitoring screen has changed abnormally based on the difference.
[0169] Step S3110: Detect whether there are any abnormal changes in the monitoring image.
[0170] If the monitoring screen changes abnormally, it means that there may be a dangerous situation that threatens the user's personal safety in the current environment. The second video image can be input into the above-mentioned abnormal image classification model so that the abnormal image classification model can identify the image type of the second video image and obtain the output of the abnormal image classification model. The output may include: ① images including flames; ② images including animals; ③ images including people, one or more combinations thereof.
[0171] If the output only includes the first case, the process may proceed to step S3120 to perform a fire identification warning.
[0172] If the output only includes the second case, the process proceeds to step S3130 to perform an animal identification warning.
[0173] If the output only includes the third case, the process proceeds to step S3140 to perform a person recognition warning.
[0174] If the output includes the two situations ① and ②, step S3120 and step S3130 can be entered at the same time to perform fire identification warning and animal identification warning at the same time.
[0175] If the output includes the two situations ① and ③, step S3120 and step S3140 can be entered at the same time to perform fire identification warning and human identification warning at the same time.
[0176] If the output includes situations ② and ③, step S3130 and step S3140 can be entered at the same time to perform animal recognition warning and human recognition warning at the same time.
[0177] If the output includes the three situations ①, ②, and ③, step S3120, step S3130, and step S3140 can be entered simultaneously to perform fire identification warning, animal identification warning, and human identification warning at the same time.
[0178] If there is no abnormal change in the monitoring image, it means that the current environment is safe. At this time, you can return to step S370 and continue to scan the environment for 1 second per frame.
[0179] Step S3120: Perform fire identification and warning.
[0180] The abnormal event is determined to be a fire, and a warning message of the fire is generated, and the process proceeds to step S3180 to perform a comprehensive information analysis based on the warning message of the fire.
[0181] Step S3130: Perform animal identification warning.
[0182] The abnormal event is determined to be an animal intrusion, and an early warning message of an animal intrusion is generated, and the process proceeds to step S3180 to perform a comprehensive information analysis based on the early warning message of an animal intrusion.
[0183] Step S3140: Perform a person recognition warning.
[0184] Perform face recognition on the second video image to obtain the identity of the person in the second video image, and enter step S3150. When the user is alone in the room and the mobile terminal or cloud only stores the user's identity information, determine whether the person is the owner based on the identified identity.
[0185] Step S3150: Detect whether the person in the second video image is the host.
[0186] When a user occupies a room alone and the mobile terminal or cloud only stores the user's identity information, the recognized person's identity is used to determine whether the person is the owner, thereby determining whether a stranger has entered the room.
[0187] If the person in the second video image is not the owner, it is determined that the person in the second video image is a stranger, and the process goes to step S3160 to perform a stranger identification warning.
[0188] If the person in the second video image is the owner, and the user is alone in the room, it means that the abnormal change in the monitoring screen is caused by the owner, and the reason for the abnormal change in the monitoring screen is most likely that the owner has fallen. At this time, step S3170 is entered to detect whether the owner has fallen.
[0189] Step S3160: Stranger identification warning.
[0190] The abnormal event is determined to be a stranger breaking in, and a warning message of a stranger breaking in is generated, and the process goes to step S3180 to perform a comprehensive information analysis based on the warning message of a stranger breaking in.
[0191] Step S3170: Detect whether the owner falls.
[0192] The body joints of the person in the second video image are obtained, and whether the person falls is detected based on the body joints.
[0193] If it is detected that the owner has fallen, the abnormal event is determined to be the owner's fall, and an early warning message of the owner's fall is generated, and the process goes to step S3180 to perform a comprehensive information analysis based on the early warning message of the owner's fall.
[0194] If it is detected that the owner has not fallen, the process returns to step S370 and continues to scan the environment for 1 second per frame.
[0195] Step S3180: Comprehensive information analysis.
[0196] Conduct comprehensive information analysis based on the received warning information of abnormal events.
[0197] If an early warning message of an abnormal event (for example, a fire, an animal breaking in, a stranger breaking in, or the owner falling) is received, the process proceeds to step S3190 to perform early warning or alarm processing based on the early warning message of the abnormal event.
[0198] If early warning information of at least two abnormal events is received, the final early warning information can be determined according to the priority of the at least two abnormal events, and step S3190 is entered to perform early warning or alarm processing according to the final early warning information. Among them, the final early warning information can be one or at least two. The priority of the abnormal events can be set in advance. For example, the priority of fire, animal intrusion, stranger intrusion or owner fall can be set in descending order. If early warning information of three abnormal events, fire, animal intrusion and owner fall, is received, the final early warning information can be the early warning information of fire, or the final early warning information can include early warning information of three abnormal events, fire, animal intrusion and owner fall.
[0199] Step S3190: Early warning and / or alarm processing.
[0200] Based on the received warning information, a warning and / or alarm process is performed. The warning and / or alarm process may include, but is not limited to, a photoelectric alarm prompt, a phone call to a family member, a text message to a family member, a direct call to the alarm number and a location message, etc. After the warning and / or alarm, the system proceeds to step S3200 to interact with the user through voice communication, further obtain the user's physical condition, or perform relevant operations according to the user's instructions (such as stopping the warning).
[0201] Step S3200: User voice interaction processing.
[0202] Interact with the user through voice, obtain the user's physical condition or perform relevant operations according to the user's instructions (such as stopping the warning), and end the abnormal monitoring after the abnormality is resolved (such as the user confirms the end of the monitoring function).
[0203] For the parts of step S310 to step S3200 that are not described in detail, please refer to the relevant parts above.
[0204] This application proposes the above-mentioned abnormal monitoring method in response to the safety needs of business travelers, providing a convenient and accessible way to meet the personal and financial safety needs of users in temporary rest places. It has made unique designs and optimizations based on the characteristics of separate travel spaces and vertical image collection of mobile terminals. Users only need to place the mobile terminal in an appropriate position, and the mobile terminal can monitor environmental changes in real time in unfamiliar environments such as hotels and inns, effectively identify dangerous situations that may occur in hotel spaces, and automatically give users voice reminders or alarms when dangerous situations (such as the above-mentioned abnormal events) occur, and take photos and videos as evidence, thereby maximizing the protection of the personal and financial safety of mobile phone users when they are away, simplifying the cost of setting up the abnormal monitoring system, and making it simple and convenient to operate.
[0205] In particular, when the mobile terminal is a foldable phone, the upright feature of the foldable phone can be used to simplify the placement of the mobile terminal, providing business and travelers with a convenient life safety monitoring and protection system, which can protect the personal safety of travelers or people living alone at a low cost, fills the application defects of foldable phones in monitoring, and can improve the user experience and user satisfaction of foldable phones.
[0206] In addition, the present application can be run on a mobile terminal or in a combination of a mobile terminal and the cloud. The advantage of local operation is that the user's privacy is guaranteed, and all data is processed and stored locally. The advantage of the combination of local and cloud is that the data is stored in the cloud, there is no fear of data loss, and the third party assists in security operations, and personal safety is more guaranteed. That is, the data such as the abnormal change screen (the first video image with abnormal changes), abnormal events and their related monitoring videos (the second video image) in the present application can be stored in the local memory of the mobile terminal, and / or encrypted and stored in the cloud memory connected to the mobile terminal after obtaining user authorization. Storing in the local memory of the mobile terminal can better protect the user's privacy. Storing in the cloud memory can avoid data loss and ensure the integrity of the data.
[0207] See also Figure 19 , Figure 19 FIG3 is a schematic diagram of the structure of an abnormality monitoring device provided in an embodiment of the present application. The abnormality monitoring device 300 is applied to a mobile terminal. The abnormality monitoring device 300 includes an image acquisition module 310, an abnormality detection module 320, an event detection module 330, and an abnormality warning module 340.
[0208] The image acquisition module 310 is configured to capture a first video image captured by a first camera and a second video image captured by a second camera.
[0209] In some embodiments, the image acquisition module 310 is also used to detect whether the mobile terminal is in a stationary state; when the mobile terminal is in a stationary state, detect whether the mobile terminal is in a monitoring position; when the mobile terminal is in the monitoring position, acquire a first video image captured by the first camera and a second video image captured by the second camera.
[0210] In some embodiments, the image acquisition module 310 is further used to detect whether the first camera and the second camera of the mobile terminal are blocked when the mobile terminal is in a stationary state; when the first camera and the second camera are not blocked, detect whether the shooting range of the first camera of the mobile terminal covers the door and window range of the room where the mobile terminal is located; when the shooting range of the first camera covers the door and window range, determine that the mobile terminal is in a monitoring position.
[0211] In some embodiments, the image acquisition module 310 is further configured to detect whether the mobile terminal is in a stationary state when the folding screen of the mobile terminal is in a semi-folded state.
[0212] In some embodiments, the image acquisition module 310 is further configured to call the first camera and the second camera in response to activation of the monitoring function.
[0213] In some embodiments, the image acquisition module 310 is further configured to display the first video image and the second video image, wherein the second video image is displayed on the first video image, and the second video image does not completely cover the first video image.
[0214] The abnormality detection module 320 is used to detect whether abnormal changes occur in the monitoring image based on multiple frames of the first video image.
[0215] In some embodiments, the anomaly detection module 320 is also used to generate baseline information based on multiple frame images before the current frame image in the first video image, and the baseline information includes a baseline image and / or a baseline feature vector; based on the current frame image and the baseline information, detect whether abnormal changes occur in the monitoring screen.
[0216] In some embodiments, the anomaly detection module 320 is further used to perform a differential operation on the baseline image in the baseline information and the current frame image to obtain a differential value; if the differential value is greater than a differential threshold, it is determined that an abnormal change has occurred in the monitoring screen; if the differential value is less than or equal to the differential threshold, it is determined that no abnormal change has occurred in the monitoring screen.
[0217] In some embodiments, the anomaly detection module 320 is also used to extract the feature vector of the current frame image; calculate the distance between the feature vector of the current frame image and the benchmark feature vector in the benchmark information; if the distance is greater than the distance threshold, determine that an abnormal change has occurred in the monitoring screen; if the distance is less than or equal to the distance threshold, determine that no abnormal change has occurred in the monitoring screen.
[0218] In some embodiments, the anomaly detection module 320 is further used to calculate an average image of multiple frame images before the current frame image as a reference image; and / or calculate an average value of feature vectors of multiple frame images before the current frame image as a reference feature vector.
[0219] The event detection module 330 is configured to detect an abnormal event based on the second video image when an abnormal change occurs in the monitoring screen.
[0220] In some embodiments, the event detection module 330 is further configured to identify the image type of the second video image when an abnormal change occurs in the monitoring image; and determine an abnormal event based on the image type.
[0221] In some embodiments, the event detection module 330 is further configured to perform face detection on the second video image if the image type is an image containing a person; and determine an abnormal event based on the face detection result.
[0222] In some embodiments, the event detection module 330 is further configured to extract human joint points in the second video image if the face detection result indicates that there are no strangers in the second video image; and determine an abnormal event based on the human joint points.
[0223] In some embodiments, the event detection module 330 is also used to determine the moving speed of the human body center of mass based on the human body joints; detect whether the user falls based on the moving speed to obtain a first detection result; detect whether the user falls based on the height of the neck joint to obtain a second detection result; detect whether the user falls based on the relative position relationship between the shoulder joint and the waist joint to obtain a third detection result; if the first detection result, the second detection result, and the third detection result all indicate that the user has fallen, it is determined that the abnormal event is a user fall.
[0224] In some embodiments, the event detection module 330 is further configured to determine that the abnormal event is an intrusion of a stranger if the face detection result indicates that the person in the second video image is a stranger.
[0225] In some embodiments, the event detection module 330 is further configured to determine that the abnormal event is a fire if the image type is an image containing flames.
[0226] In some embodiments, the event detection module 330 is further configured to determine that the abnormal event is an animal intrusion if the image type is an image containing an animal.
[0227] The abnormality warning module 340 is used to issue a warning message according to the abnormal event.
[0228] Those skilled in the art can clearly understand that the above devices provided in the embodiments of the present application can implement the methods provided in the embodiments of the present application. The specific working processes of the above-described devices and modules can refer to the corresponding processes of the methods in the embodiments of the present application, which will not be repeated here.
[0229] In the embodiments provided in the present application, the coupling, direct coupling or communication connection between the modules shown or discussed may be an indirect coupling or communication coupling through some interfaces, devices or modules, and may be electrical, mechanical or other forms, and the embodiments of the present application do not impose specific limitations on this.
[0230] In addition, the functional modules in the embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0231] See also Figure 20 , Figure 20 : This is a schematic diagram of the structure of a mobile terminal provided in one embodiment of the present application. Mobile terminal 400 may include a first camera 410, a second camera 420, a memory 430, and a processor 440. The memory 430 stores an application program, and processor 440 executes the method provided in the embodiment of the present application when invoking the application program. Mobile terminal 400 may be the same as mobile terminal 100 described above. First camera 410 may be the same as first camera 131 described above. Second camera 420 may be the same as second camera 132 described above.
[0232] The processor 440 may include one or more processing cores. The processor 440 uses various interfaces and lines to connect various components within the mobile terminal 400 and is used to run or execute instructions, programs, code sets, or instruction sets stored in the memory 430, as well as call and execute data stored in the memory 430, perform various functions of the mobile terminal 400, and process data.
[0233] The processor 440 can be implemented in at least one hardware form of digital signal processing (DSP), field programmable gate array (FPGA), and programmable logic array (PLA). The processor 440 can integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem is used to handle wireless communications. It is understandable that the above-mentioned modem may not be integrated into the processor 440, but may be implemented separately through a communication chip.
[0234] The memory 430 may include a random access memory (RAM) or a read-only memory (ROM). The memory 430 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 430 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function, instructions for implementing the above-mentioned various method embodiments, and the like. The data storage area may store data created by the mobile terminal 400 during use, and the like.
[0235] An embodiment of the present application provides a computer-readable storage medium having program code stored thereon. When a processor calls the program code, the method provided in the embodiment of the present application is executed.
[0236] The computer-readable storage medium may be an electronic memory such as a flash memory, an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a hard disk, or a ROM.
[0237] In some embodiments, the computer-readable storage medium includes a non-volatile computer-readable medium (Non-Transitory Computer-Readable Storage Medium, referred to as Non-TCRSM). The computer-readable storage medium has storage space for program codes that execute any method step in the above method. These program codes can be read from or written into one or more computer program products. The program code can be compressed in an appropriate form.
[0238] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for monitoring abnormalities, characterized in that: Applied to a mobile terminal, the mobile terminal includes a first camera and a second camera, the viewing angle of the first camera is larger than that of the second camera, and the method includes: Acquire a first video image captured by the first camera and a second video image captured by the second camera; detecting, based on multiple frames of the first video image, whether abnormal changes occur in the monitoring image; When an abnormal change occurs in the monitoring screen, detecting an abnormal event based on the second video image; According to the abnormal event, an early warning message is issued.
2. The method according to claim 1, characterized in that The detecting whether abnormal changes occur in the monitoring screen according to the multiple frames of the first video image includes: generating reference information according to a plurality of frames of images preceding a current frame of image in the first video image, the reference information including a reference image and / or a reference feature vector; Based on the current frame image and the reference information, it is detected whether an abnormal change occurs in the monitoring screen.
3. The method according to claim 2, characterized in that The detecting whether an abnormal change occurs in the monitoring screen according to the current frame image and the reference information includes: Performing a differential operation on the reference image in the reference information and the current frame image to obtain a differential value; If the difference value is greater than the difference threshold, it is determined that an abnormal change occurs in the monitoring image; If the difference value is less than or equal to the difference threshold, it is determined that no abnormal change occurs in the monitoring image.
4. The method according to claim 2, characterized in that The detecting whether an abnormal change occurs in the monitoring screen according to the current frame image and the reference information includes: Extracting a feature vector of the current frame image; Calculating the distance between the feature vector of the current frame image and the reference feature vector in the reference information; If the distance is greater than the distance threshold, it is determined that an abnormal change occurs in the monitoring image; If the distance is less than or equal to the distance threshold, it is determined that no abnormal changes occur in the monitoring image.
5. The method according to claim 2, characterized in that The generating of the reference information according to the plurality of frames of images preceding the current frame of the first video image includes: Calculating an average image of multiple frames of images before the current frame of image as a reference image; and / or An average value of the feature vectors of multiple frames of images before the current frame of image is calculated as a reference feature vector.
6. The method according to any one of claims 1 to 5, characterized in that When an abnormal change occurs in the monitoring screen, detecting an abnormal event according to the second video image includes: When an abnormal change occurs in the monitoring image, identifying the image type of the second video image; An abnormal event is determined based on the image type.
7. The method according to claim 6, characterized in that The determining of an abnormal event according to the image type includes: If the image type is an image containing a person, performing face detection on the second video image; Determine abnormal events based on face detection results.
8. The method according to claim 7, characterized in that Determining an abnormal event based on the face detection result includes: If the face detection result indicates that there are no strangers in the second video image, extracting human body joints in the second video image; Abnormal events are determined according to the human body joint points.
9. The method according to claim 8, characterized in that The human body joints include at least a neck joint, a shoulder joint, and a waist joint. The determining of abnormal events based on the human body joints includes: Determining the moving speed of the center of mass of the human body according to the human body joint points; detecting whether the user falls according to the movement speed to obtain a first detection result; detecting whether the user has fallen according to the height of the neck joint point to obtain a second detection result; detecting whether the user has fallen according to the relative positional relationship between the shoulder joint point and the waist joint point to obtain a third detection result; If the first detection result, the second detection result, and the third detection result all indicate that the user has fallen, it is determined that the abnormal event is a fall of the user.
10. The method according to claim 7, characterized in that Determining an abnormal event based on the face detection result includes: If the face detection result indicates that there is a stranger in the second video image, the abnormal event is determined to be a stranger breaking in.
11. The method according to claim 6, characterized in that The determining of an abnormal event according to the image type includes: If the image type is an image containing flames, it is determined that the abnormal event is the occurrence of a fire.
12. The method according to claim 6, characterized in that The determining of an abnormal event according to the image type includes: If the image type is an image containing animals, it is determined that the abnormal event is an animal intrusion.
13. The method according to claim 1, wherein The acquiring of a first video image captured by the first camera and a second video image captured by the second camera includes: Detecting whether the mobile terminal is in a stationary state; When the mobile terminal is in a stationary state, detecting whether the mobile terminal is in a monitoring position; When the mobile terminal is at the monitoring position, a first video image captured by the first camera and a second video image captured by the second camera are acquired.
14. The method according to claim 13, characterized in that When the mobile terminal is in a stationary state, detecting whether the mobile terminal is in a monitoring position includes: When the mobile terminal is in a stationary state, detecting whether the first camera and the second camera of the mobile terminal are blocked; When the first camera and the second camera are not blocked, detecting whether the shooting range of the first camera of the mobile terminal covers the door and window range of the room where the mobile terminal is located; When the shooting range of the first camera covers the door and window range, it is determined that the mobile terminal is in the monitoring position.
15. The method according to claim 13, characterized in that The mobile terminal is a foldable screen mobile terminal, and detecting whether the mobile terminal is in a stationary state includes: When the folding screen of the mobile terminal is in a semi-folded state, it is detected whether the mobile terminal is in a stationary state.
16. The method according to any one of claims 13 to 15, characterized in that: Before acquiring the first video image captured by the first camera and the second video image captured by the second camera, the method further includes: In response to the monitoring function being turned on, the first camera and the second camera are called.
17. The method according to claim 1, wherein After acquiring the first video image captured by the first camera and the second video image captured by the second camera, the method further includes: The first video image and the second video image are displayed, wherein the second video image is displayed on the first video image and the second video image does not completely cover the first video image.
18. An abnormality monitoring device, characterized in that: Applied to a mobile terminal, the mobile terminal includes a first camera and a second camera, the viewing angle of the first camera is larger than that of the second camera, and the device includes: An image acquisition module, configured to capture a first video image captured by a first camera and a second video image captured by a second camera; an abnormality detection module, configured to detect whether abnormal changes occur in the monitoring image based on multiple frames of the first video image; an event detection module, configured to detect an abnormal event based on the second video image when an abnormal change occurs in the monitoring screen; The abnormal warning module is used to issue warning information according to the abnormal event.
19. A mobile terminal, characterized in that: include: A first camera, a second camera, a memory, and a processor, wherein the first camera has a larger viewing angle than the second camera, an application is stored in the memory, and the processor executes the method according to any one of claims 1 to 17 when calling the application.
20. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, and when the processor calls the program code, the method according to any one of claims 1 to 17 is executed.