Pet state recognition method and apparatus, electronic device, and computer readable medium
By collecting location signals and monitoring audio and video through pet identification devices, the type and status of pets can be determined, solving the problem of pet status identification in complex spatial environments, achieving effective and accurate identification of pet status, and reducing the risk of abnormal pet behavior.
Patent Information
- Application Number
- CN202411942775.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In complex environments, it is difficult to effectively and accurately monitor the real-time status of pets, leading to frequent cases of pet deaths due to abnormal behavior.
By receiving location signals through pet identification devices, determining the location, and collecting video and audio data in real time, the system combines the signal source location with real-time monitoring audio and video data to determine the pet type, collect monitoring audio and video data, and determine pet status description information, thereby providing pet behavior description information and health description information.
It enables effective and accurate identification of pet status in complex spatial environments, improving the accuracy and efficiency of pet status identification and reducing the risk of abnormal pet behavior.
Smart Images

Figure CN119380417B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the technical field of computer technology, and in particular, to a pet state recognition method and device, an electronic device, and a computer readable medium. BACKGROUND
[0002] A pet refers to a general term for animals that are raised and cared for by humans. As pets have attributes such as companionship, more and more people begin to raise pets as family members. However, as common pets such as cats and dogs have characteristics such as small size and high flexibility of movement, it is difficult to effectively monitor the real-time state of pets in indoor spaces with large areas and complex spatial environments. According to related reports, there are cases of pet deaths caused by abnormal behaviors such as accidental ingestion. Therefore, how to effectively and accurately recognize the state of a pet has become a problem to be solved.
[0003] The above information disclosed in this BACKGROUND section is only for the purpose of enhancing the understanding of the background of the present inventive concepts, and therefore, it can contain information that does not form the prior art that is already known to those of ordinary skill in the art. SUMMARY
[0004] The summary of the present disclosure is intended to introduce the concepts in a simplified form, which will be described in detail in the detailed description section below. The summary of the present disclosure is not intended to identify key or essential features of the claimed technology nor is it intended to be used to limit the scope of the claimed technology.
[0005] Some embodiments of the present disclosure provide a pet state recognition method, device, electronic device, and computer readable medium to solve one or more of the technical problems mentioned in the background section.
[0006] In a first aspect, some embodiments of the present disclosure provide a pet state recognition method applied to a pet recognition device, including: in response to receiving at least one positioning signal sent by the pet recognition device, determining a signal source position according to the at least one positioning signal; collecting real-time monitoring audio and video according to the signal source position; determining a pet type according to the signal source position and monitoring video included in the monitoring audio and video; in response to the pet type and audio description information satisfying a first condition, determining pet state description information according to the signal source position and the monitoring audio and video, wherein the audio description information is an audio description of monitoring audio included in the monitoring audio and video, and the pet state description information includes pet behavior description information and pet health description information; and in response to the pet type and the audio description information not satisfying the first condition, determining the pet state description information according to the signal source position and monitoring video included in the monitoring audio and video.
[0007] Optionally, the determining the pet type according to the signal source position and the monitoring video included in the monitoring audio and video comprises: determining the pet type according to the signal source position, the monitoring video included in the monitoring audio and video, and a pet identity corresponding to the positioning signal, and binding the pet identity with the recognized pet target.
[0008] Optionally, the determining the signal source position according to the at least one positioning signal comprises: for each positioning signal in the at least one positioning signal, performing the following processing steps: determining a target information set corresponding to the positioning signal, wherein the target information in the target information set comprises: a signal time difference, a first coordinate, and a second coordinate, the signal time difference representing a time difference between the first coordinate corresponding signal receiver and the second coordinate corresponding signal receiver receiving the positioning signal, determining a candidate signal source position according to the target information set; and performing signal source position correction according to the obtained candidate signal source position set to generate the signal source position.
[0009] Optionally, the pet behavior description information comprises: a pet behavior type and a pet behavior duration, wherein the pet behavior duration represents a duration for which the pet maintains the behavior corresponding to the pet behavior type, the pet health description information comprises: a pet health degree, and the method further comprises: updating the pet health degree according to the pet behavior type and the pet behavior duration to obtain an updated pet health degree; in response to the updated pet health degree being located in a first health degree interval, sending a health degree abnormality prompt information to the target client; and in response to the pet behavior type being an abnormal behavior type and the pet behavior duration being greater than an abnormal behavior duration threshold corresponding to the pet behavior type, sending a pet behavior abnormality prompt information to the target client.
[0010] Optionally, the determining the pet type according to the signal source position and the monitoring video included in the monitoring audio and video comprises: performing shallow image feature extraction on the monitoring video to obtain a shallow image feature sequence; determining a signal source type according to the shallow image feature sequence, wherein the signal source type comprises: a living body type and a non-living body type, and the living body type comprises: a pet living body type and a non-pet living body type; in response to the signal source type being the pet living body type, determining a region of interest according to the signal source position, wherein the region of interest comprises: at least one superimposed region of interest; performing feature windowing on each shallow image feature in the shallow image feature sequence according to the at least one superimposed region of interest to generate a windowed shallow image feature, thereby obtaining a windowed shallow image feature sequence; performing deep image feature extraction on the windowed shallow image feature sequence to generate a deep image feature, thereby obtaining a deep image feature sequence; and determining the pet type according to the deep image feature sequence.
[0011] Optionally, the audio description information comprises audio source description information and audio quality description information, the audio source description information represents a sound source included in the monitoring audio, and the audio description information is generated by: performing audio feature extraction on the monitoring audio to generate audio features; generating the audio source description information by using an audio source recognition model and the audio features; and generating the audio quality description information according to an audio quality scoring model, the audio features and the audio source description information.
[0012] Optionally, the pet state description information is determined according to the signal source position and the monitoring audio and video, comprising: performing video feature extraction on the monitoring video to generate video features, wherein the video feature extraction model shares a shallow image feature extraction model and a deep image feature extraction model included in the pet recognition model; performing audio feature extraction on the monitoring audio to generate audio features, wherein the audio feature extraction model is an audio feature extraction model in the audio description information generation model; performing feature fusion on the video features and the audio features to generate fused features; and generating the pet state description information according to the fused features and a pet state description information predictor, wherein the pet state description information predictor comprises a pet behavior type predictor and a pet health degree predictor.
[0013] Optionally, the method further comprises: determining whether to allow video uploading; in response to allowing, performing video local blurring on the monitoring video to obtain a local-blurred monitoring video; performing video compression on the local-blurred monitoring video to generate a compressed video; and respectively storing the compressed video at a local end and a cloud end.
[0014] Optionally, the monitoring video is locally blurred to obtain a local-blurred monitoring video, comprising: determining a region confidence of each superimposed region of interest in the at least one superimposed region of interest, wherein the region confidence represents a confidence that the pet is included in the corresponding superimposed region of interest; selecting a superimposed region of interest that satisfies a second condition from the at least one superimposed region of interest as a target region of interest according to a region size of the superimposed region of interest and a region confidence corresponding to the superimposed region of interest; performing inverse selection cropping on each shallow image feature in the shallow image feature sequence according to the target region of interest to generate a cropped shallow image feature, to obtain a cropped shallow image feature sequence; determining a position of a non-pet living body type object in the monitoring video according to the cropped shallow image feature sequence and a non-pet living body positioning model, to obtain a non-pet living body position; and performing video local blurring at the non-pet living body position in the monitoring video to obtain a local-blurred monitoring video.
[0015] In a second aspect, some embodiments of the present disclosure provide a pet identification device, comprising: a pet locator, wherein the pet locator comprises: at least two ultra-wideband signal transmitters, the pet locator being fixed to a surface of a pet in a wearable form; a signal collection device, wherein the signal collection device comprises: at least three ultra-wideband signal receivers, a video collection device, and an audio collection device, the at least three ultra-wideband signal receivers, the video collection device, and the audio collection device being directed to the same collection surface; one or more processors; and a storage device having one or more programs stored thereon, which when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect.
[0016] In a third aspect, some embodiments of the present disclosure provide a computer readable medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect.
[0017] The above various embodiments of the present disclosure have the following beneficial effects: through the pet state recognition method of some embodiments of the present disclosure, effective and accurate pet state recognition is achieved. Specifically, the reason why effective and accurate pet state recognition cannot be achieved is that common pets such as cats and dogs have the characteristics of small size and high flexibility, and it is difficult to effectively monitor and recognize the real-time state of the pet in an indoor space with a complex space environment (such as an indoor space with furniture). Based on this, the present disclosure first, in response to receiving at least one positioning signal sent by the pet recognition device, determines the signal source position according to the at least one positioning signal, thereby obtaining the real-time position of the signal source. Second, according to the signal source position, real-time monitoring audio and video are collected, then according to the signal source position and the monitoring video included in the monitoring audio and video, the pet type is determined. In this way, the accurate pet type is obtained. Further, in response to the pet type and the audio description information satisfying a first condition, the pet state description information is determined according to the signal source position and the monitoring audio and video, wherein the audio description information is an audio description of the monitoring audio included in the monitoring audio and video, and the pet state description information includes pet behavior description information and pet health description information. Finally, in response to the pet type and the audio description information not satisfying the first condition, the pet state description information is determined according to the signal source position and the monitoring video included in the monitoring audio and video. In practice, the sound of some pets is small or the noise is large when collecting, so that the quality of the obtained monitoring audio is poor. At this time, the monitoring audio is used as one of the bases for determining the pet state description information, which will increase the noise and affect the accuracy of the obtained pet state description information. Therefore, when the monitoring audio is poor, only the video and the signal source position are combined to determine the pet behavior. In this way, effective and accurate pet state recognition is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0018] The above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent upon consideration of the following detailed description, taken in conjunction with the accompanying drawings. Throughout the drawings, like or similar reference numerals designate identical or similar elements throughout the several views. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.
[0019] Figure 1 is a flowchart of some embodiments of the pet state recognition method according to the present disclosure;
[0020] Figure 2 is a front view of a signal collection device included in pet recognition equipment;
[0021] Figure 3 is a schematic view of a pet locator included in pet recognition equipment;
[0022] Figure 4is a schematic diagram of a single-pet target recognition process;
[0023] Figure 5 is a schematic diagram of a multi-pet target recognition process;
[0024] Figure 6 is another schematic diagram of a multi-pet target recognition process;
[0025] Figure 7 is another flowchart of some embodiments of the pet status recognition method according to the present disclosure;
[0026] Figure 8 is a structural schematic diagram of an electronic device suitable for being used to implement some embodiments of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein. Rather, these embodiments are provided so as to more completely and comprehensively understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are merely for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.
[0028] It should also be noted that, for ease of description, only parts related to the present application are shown in the drawings. The embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0029] It should be noted that the concepts of "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0030] It should be noted that the adjectives "one", "multiple" mentioned in the present disclosure are illustrative and not limiting, and those skilled in the art should understand that, unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0031] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0032] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0033] Reference Figure 1 is a flowchart 100 of some embodiments of the pet status recognition method according to the present disclosure. The pet status recognition method comprises the following steps:
[0034] Step 101, in response to receiving at least one positioning signal sent by the pet identification device, determining the signal source position according to the at least one positioning signal.
[0035] In some embodiments, the execution subject of the pet status identification method (e.g., a computing device) can determine the signal source position according to the at least one positioning signal in response to receiving at least one positioning signal sent by the pet identification device. The at least one positioning signal can be emitted by the pet locator included in the pet identification device. The signal source position can represent the position of the pet locator on the pet. In practice, since the pet locator is fixed to the pet's body surface in a wearable form, the signal source position can represent the position of the pet. The above-mentioned execution subject can accept the at least one positioning signal through the ultra-wideband signal receiver included in the pet identification device.
[0036] As an example, the above-mentioned execution subject can determine the signal source position according to the incident angle of the positioning signal accepted by the ultra-wideband signal receiver and the signal propagation time. Specifically, since the speed of light is known and the receiver positions of the at least three ultra-wideband signal receivers are known, the distance from the ultra-wideband signal receiver to the signal source position can be obtained according to the signal propagation time and the speed of light. Thus, the signal source position can be calculated through the incident angle of the positioning signal to the ultra-wideband signal receiver.
[0037] As another example, the signal acquisition device can include at least three ultra-wideband signal receivers, a video acquisition device, and an audio acquisition device, and the above-mentioned at least three ultra-wideband signal receivers, the above-mentioned video acquisition device, and the above-mentioned audio acquisition device are directed to the same acquisition surface. Specifically, referring to Figure 2 The front view of the signal acquisition device included in the pet identification equipment is shown, wherein the signal acquisition device includes at least three ultra-wideband signal receivers 1 (specifically, Figure 2 Two groups of four ultra-wideband signal receivers 1, a video acquisition device 2, and an audio acquisition device 3 are shown, and the above-mentioned at least three ultra-wideband signal receivers 1, the above-mentioned video acquisition device 2, and the above-mentioned audio acquisition device 3 are directed to the same acquisition surface. Specifically, the video acquisition device includes a main camera 5 and a sub-camera 6. In practice, the main camera 5 is a wide-angle camera. The sub-camera 6 is an infrared camera. The audio acquisition device 3 is a directional microphone. In addition, the signal acquisition device also includes a rotating table 4. The signal acquisition device can be installed on the top surface of the wall through the rotating table 4, and the change of the acquisition surface of the signal acquisition device can be controlled by controlling the rotation of the rotating table 4. In practice, the rotating table 4 supports horizontal rotation and 30-degree pitch rotation.
[0038] As another example, referring to Figure 3The pet identification equipment shown includes a pet locator. The pet locator can be a pet collar 7. The pet collar 7 includes at least two ultra-wideband signal transmitters 8. Figure 3 The two ultra-wideband signal transmitters 8 are arranged in a symmetrical structure. For three ultra-wideband signal transmitters 7, one ultra-wideband signal transmitter 7 can be arranged on the pet collar 7 every 120 degrees. Specifically, the ultra-wideband signal transmitters 7 have low power consumption characteristics, and therefore are powered by a button cell. In addition, the pet locator corresponds to a pet identity, i.e., the at least two ultra-wideband signal transmitters correspond to the same pet identity. The pet identity is a unique identity of the pet. In addition, the ultra-wideband signal transmitters correspond to a device identifier. For example, when the pet locator includes two ultra-wideband signal transmitters, i.e., an ultra-wideband signal transmitter A and an ultra-wideband signal transmitter B. The ultra-wideband signal transmitter A corresponds to a device identifier A. The ultra-wideband signal transmitter B corresponds to a device identifier B. The device identifier corresponding to each ultra-wideband signal transmitter is encoded in the frame structure of the ultra-wideband signal. Alternatively, the pet locator can also be fixed on pet clothes, pet tags, etc. For example, magic glue or the like can be used to fix the pet locator on the pet clothes. Alternatively, the pet locator can also be fixed in the pet tag by embedding.
[0039] It should be noted that the above computing device can be hardware or software. When the computing device is hardware, it can be implemented as a single terminal device. When the computing device is software, it can be installed in the above-mentioned hardware devices. It can be implemented as a single software or software module. Herein, no specific limitation is made. In practice, the above computing device can be a processor included in the signal acquisition apparatus, such as a CPU (Central Processing Unit, Central Processing Unit).
[0040] Step 102, real-time acquisition of monitoring audio and video according to the signal source position.
[0041] In some embodiments, the above execution subject can real-time acquisition of monitoring audio and video according to the signal source position. The monitoring audio and video can be a video containing video pictures and audio at the signal source position. In practice, the above execution subject can control the video acquisition device and the audio acquisition device to acquire the video pictures and audio at the signal source position as the above monitoring audio and video.
[0042] Step 103, determining the pet type according to the signal source position and the monitoring video included in the monitoring audio and video.
[0043] In some embodiments, the execution subject can determine the pet type according to the signal source position and the monitoring video included in the monitoring audio-video. The pet type can represent the specific type of the pet at the signal source position. For example, the pet type can include Maine cat type, cow cat type, three-flower cat type, teddy dog type, and golden retriever type. In practice, the execution subject can perform picture recognition on the picture at the signal source position in the monitoring video to determine the pet type.
[0044] In some optional implementations of some embodiments, the execution subject determines the pet type according to the signal source position and the monitoring video included in the monitoring audio-video, which includes:
[0045] According to the signal source position, the monitoring video included in the monitoring audio-video, and the pet identity corresponding to the positioning signal, the pet type is determined, and the pet identity is bound to the recognized pet target.
[0046] In practice, the execution subject can determine the region of interest in the monitoring video according to the signal source position. The region of interest can be a region containing a pet. Then, pet recognition is performed in the region of interest to recognize the pet target and the pet type corresponding to the pet target. Then, the execution subject can assign the pet identity corresponding to the positioning signal to the pet target to achieve binding and automatic calibration.
[0047] By combining the pet identity with the visual recognition algorithm, not only is the automatic calibration of the pet target achieved, replacing the manual calibration method and improving the calibration efficiency. At the same time, the accuracy and efficiency of pet recognition are significantly improved. In addition, the training samples obtained through automatic calibration are used to continuously train and optimize the model to further improve the recognition accuracy and efficiency of the model, while achieving the effect of model customization.
[0048] As an example, refer to Figure 4 The single-pet target recognition process diagram shown in FIG. 1, in which pet A (e.g., dog A) emits a positioning signal to the signal acquisition device through a pet locator worn by pet A. At this time, the signal acquisition device can acquire a monitoring video. The region of interest in the monitoring video is a region containing pet A. Then, pet recognition can be performed in the region of interest to recognize pet target A, and the pet identity corresponding to pet A is assigned to the recognized pet target A to achieve binding and automatic calibration.
[0049] As another example, refer to Figure 5The shown is a schematic diagram of a multi-pet target recognition process, wherein pet A1 (e.g., dog A1) and pet A2 (e.g., dog A2) emit positioning signals to the signal acquisition device through the pet locator worn. At this time, the signal acquisition device can collect monitoring video. The monitoring video can be two different monitoring videos for pet A1 and pet A2, or a monitoring video containing pet A1 and pet A2. For the former, the two monitoring videos can each include a region of interest corresponding to pet A1 and pet A2, respectively. For the latter, the monitoring video can include two regions of interest corresponding to pet A1 and pet A2, in which case there can be an intersection between the two regions of interest. Then, pet recognition can be performed in the region of interest to identify pet target A1 and assign the pet identity corresponding to pet A1 to the identified pet target A1, and identify pet target A2 and assign the pet identity corresponding to pet A2 to the identified pet target A1, to achieve binding and automatic calibration.
[0050] As a further example, see Figure 6 The shown is a schematic diagram of a multi-pet target recognition process, wherein pet A1 (e.g., dog A1) and pet B1 (e.g., cat B1) emit positioning signals to the signal acquisition device through the pet locator worn. At this time, the signal acquisition device can collect monitoring video. The monitoring video can be two different monitoring videos for pet A1 and pet B1, or a monitoring video containing pet A1 and pet B1. For the former, the two monitoring videos can each include a region of interest corresponding to pet A1 and pet B1, respectively. For the latter, the monitoring video can include two regions of interest corresponding to pet A1 and pet B1, in which case there can be an intersection between the two regions of interest. Then, pet recognition can be performed in the region of interest to identify pet target A1 and assign the pet identity corresponding to pet A1 to the identified pet target A1, and identify pet target A2 and assign the pet identity corresponding to pet B1 to the identified pet target B1, to achieve binding and automatic calibration.
[0051] Step 104, in response to the pet type and audio description information satisfying the first condition, determining the pet state description information according to the signal source position and the monitoring audio and video.
[0052] In some embodiments, in response to the pet type and the audio description information satisfying a first condition, the execution subject can determine pet state description information according to the signal source position and the monitoring audio and video. The audio description information is an audio description of the monitoring audio included in the monitoring audio and video. The pet state description information includes pet behavior description information and pet health description information. The first condition is that the average pet volume of the pet represented by the pet type is greater than or equal to a preset pet volume, and the audio quality represented by the audio description information is greater than a preset audio quality. In practice, some small animals have a small volume when they make a sound or do not make a sound, so it is not possible to effectively combine the corresponding monitoring audio to generate pet state description information. In addition, the collected monitoring audio can include the sound of multiple sources. When there are multiple sources, there can be mixing between the sounds, affecting the audio quality.
[0053] As an example, the execution subject can use a multi-modal model to obtain pet state description information by taking the picture of the signal source position in the monitoring video and the audio emitted by the signal source position in the monitoring audio as input. For example, the multi-modal model can be a Video-LLaMA2 model.
[0054] Optionally, the audio description information can include audio source description information and audio quality description information. The audio source description information represents the sound source included in the monitoring audio. The audio quality description information can represent the audio quality of the monitoring audio. The pet behavior description information is used to describe the pet behavior. The pet health description information is used to describe the current health status of the pet.
[0055] Optionally, the audio description information is generated by the following steps:
[0056] First, audio features are extracted from the monitoring audio to generate audio features.
[0057] In practice, the monitoring audio has a time sequence feature, so the audio feature extraction model uses a recurrent neural network model as the backbone network.
[0058] Second, the audio source description information is generated by an audio source identification model and the audio features.
[0059] In practice, the audio source identification model uses a fully connected layer as a predictor to predict the number of sound sources included in the monitoring audio as the audio source description information.
[0060] Third, the audio quality description information is generated according to an audio quality scoring model, the audio features, and the audio source description information.
[0061] The audio feature extraction model, the audio source identification model, and the audio quality scoring model are included in an audio description information generation model. In practice, the audio quality scoring model takes the audio feature and the audio description information as input to obtain an audio quality score for the monitoring audio. Specifically, a regression model can be used as the audio quality scoring model to obtain the audio quality description information.
[0062] At step 105, in response to the pet type and the audio description information not satisfying the first condition, pet state description information is determined according to the signal source position and the monitoring video included in the monitoring audio-video.
[0063] In some embodiments, in response to the pet type and the audio description information not satisfying the first condition, the execution subject can determine the pet state description information according to the signal source position and the monitoring video included in the monitoring audio-video.
[0064] As an example, the execution subject can input the frame of the signal source position in the monitoring video into a single-modal (video modal) model such as a YOLO (You Only Look Once) model to obtain the pet state description information.
[0065] Optionally, the present disclosure can adopt a "cloud-edge architecture", that is, the algorithms corresponding to steps 101-105 of the present disclosure can be deployed to a local device (e.g., a camera or a local server). The models involved in the present disclosure can be trained and updated in a cloud device (e.g., a cluster server), and the storage and analysis of big data can be performed. Through the "cloud-edge architecture", the computing power of the local device at the edge can be fully utilized, and the cloud device can be efficiently used for rapid model training and updating.
[0066] In addition, it should be noted that for the collection, storage, and analysis of user-related information, the corresponding subject of the user-related information has been informed of the obligation and has obtained the authorization and consent of the corresponding subject of the user-related information before performing the corresponding operation.
[0067] The above various embodiments of the present disclosure have the following beneficial effects: through the pet state recognition method of some embodiments of the present disclosure, effective and accurate pet state recognition is achieved. Specifically, the reason why effective and accurate pet state recognition cannot be achieved is that common pets such as cats and dogs have small size, high flexibility and other characteristics, and it is difficult to effectively monitor and identify the real-time state of the pet in an indoor space with a complex space environment (such as an indoor space with furniture). Based on this, the present disclosure first, in response to receiving at least one positioning signal sent by the pet recognition device, determines the signal source position according to the at least one positioning signal, thereby obtaining the real-time position of the signal source. Second, according to the signal source position, real-time monitoring audio and video are collected, then according to the signal source position and the monitoring video included in the monitoring audio and video, the pet type is determined. In this way, the accurate pet type is obtained. Further, in response to the pet type and the audio description information satisfying a first condition, the pet state description information is determined according to the signal source position and the monitoring audio and video, wherein the audio description information is an audio description of the monitoring audio included in the monitoring audio and video, and the pet state description information includes pet behavior description information and pet health description information. Finally, in response to the pet type and the audio description information not satisfying the first condition, the pet state description information is determined according to the signal source position and the monitoring video included in the monitoring audio and video. In practice, the sound of some pets is small or the noise is large when collecting, so that the quality of the obtained monitoring audio is poor. At this time, the monitoring audio is used as one of the bases for determining the pet state description information, which will increase the noise and affect the accuracy of the obtained pet state description information. Therefore, when the monitoring audio is poor, only the video and the signal source position are combined to determine the pet behavior. In this way, effective and accurate pet state recognition is achieved.
[0068] Further reference is made to Figure 7 which shows the flow 700 of some other embodiments of a pet state recognition method. The flow 700 of the pet state recognition method includes the following steps:
[0069] Step 701, in response to receiving at least one positioning signal sent by a pet recognition device, determining a signal source position according to the at least one positioning signal.
[0070] In some embodiments, the pet state recognition method executed by the execution subject (such as a computing device) in response to receiving at least one positioning signal sent by a pet recognition device, according to the at least one positioning signal, can include the following steps:
[0071] First, for each positioning signal in the at least one positioning signal, the following processing steps are performed:
[0072] A first sub-step of determining a target information set corresponding to the positioning signal.
[0073] The target information in the target information set includes a signal time difference, a first coordinate, and a second coordinate. The signal time difference represents a time difference between the first coordinate corresponding signal receiver and the second coordinate corresponding signal receiver receiving the positioning signal.
[0074] As an example, the at least three ultra-wideband signal receivers can include an ultra-wideband signal receiver A, an ultra-wideband signal receiver B, and an ultra-wideband signal receiver C. The first coordinate can be the position coordinate of the ultra-wideband signal receiver A. The second coordinate can be the position coordinate of the ultra-wideband signal receiver B. The signal time difference can be the time difference between the ultra-wideband signal receiver A and the ultra-wideband signal receiver B receiving the positioning signal.
[0075] A second sub-step of determining a candidate signal source position according to the target information set.
[0076] As an example, the at least one positioning signal includes a positioning signal A1 and a positioning signal A2. The at least two ultra-wideband signal transmitters include an ultra-wideband signal transmitter B1 and an ultra-wideband signal transmitter B2. The at least three ultra-wideband signal receivers include an ultra-wideband signal receiver C1, an ultra-wideband signal receiver C2, and an ultra-wideband signal receiver C3. The positioning signal A1 is emitted by the ultra-wideband signal transmitter B1. The positioning signal A2 is emitted by the ultra-wideband signal transmitter B2. Since the speed of light and the receiver position coordinates of the ultra-wideband signal receiver C1, the ultra-wideband signal receiver C2, and the ultra-wideband signal receiver C3 are known, for any transmitter position coordinate of the ultra-wideband signal transmitter, three sets of equations can be constructed. Taking the ultra-wideband signal receiver C1 and the ultra-wideband signal receiver C2 as an example, the transmitter position coordinate of the ultra-wideband signal receiver C1 (the first coordinate) and the transmitter position coordinate of the ultra-wideband signal receiver C2 (the second coordinate) are known, so the position difference between the first coordinate and the signal source coordinate - the position difference between the second coordinate and the signal source coordinate = the speed of light × the signal time difference. By constructing at least three sets of equations, the candidate signal source position can be solved.
[0077] A second step of correcting the signal source position according to the obtained candidate signal source position set to generate the signal source position.
[0078] In practice, since at least one positioning signal is received, multiple candidate signal source positions can be solved. In order to eliminate positioning errors, the position mean of the obtained candidate signal source position set can be used as the signal source position.
[0079] Step 702 of collecting monitoring audio and video in real time according to the signal source position.
[0080] In some embodiments, the implementation of step 702 and the technical effects brought by it can refer to the implementation of step 102 in the corresponding embodiments, which will not be described here again. Figure 1 The implementation of step 102 in the corresponding embodiments will not be described here again.
[0081] Step 703, determining the pet type according to the signal source position and the monitoring video included in the monitoring audio-video.
[0082] In some embodiments, the execution subject determines the pet type according to the signal source position and the monitoring video included in the monitoring audio-video, which can include the following steps:
[0083] Firstly, shallow image feature extraction is performed on the monitoring video to obtain a shallow image feature sequence.
[0084] Since the monitoring video is a collection of multiple monitoring pictures, the execution subject can perform shallow image feature extraction on the monitoring pictures in the monitoring video to obtain a shallow image feature sequence. The shallow image feature is an image feature of the monitoring picture under a high receptive field.
[0085] In practice, the execution subject can perform shallow image feature extraction on the monitoring video through a shallow image feature extraction model to obtain a shallow image feature sequence.
[0086] As an example, the shallow image feature extraction model includes 18 serially connected convolutional layers. Specifically, the shallow image feature extraction model includes: convolutional layer A1, convolutional layer A2, convolutional layer A3, convolutional layer A4, convolutional layer A5, convolutional layer A6, convolutional layer A7, convolutional layer A8, convolutional layer A9, convolutional layer A10, convolutional layer A11, convolutional layer A12, convolutional layer A13, convolutional layer A14, convolutional layer A15, convolutional layer A16, convolutional layer A17, and convolutional layer A18. The output of convolutional layer A1 is the input of convolutional layer A2 and the input of convolutional layer A4. Specifically, the output of convolutional layer Ai is the input of convolutional layer Ai+1 and the input of convolutional layer Ai+3. Wherein, i=2n+1, n is 0 or a positive integer, and n≤7.
[0087] Secondly, the signal source type is determined according to the shallow image feature sequence.
[0088] The signal source type includes: living body type and non-living body type. The living body type includes: pet living body type and non-pet living body type. The non-pet living body type can represent a person. In practice, the execution subject can determine the signal source type according to the shallow image feature sequence through a signal source type classifier.
[0089] As an example, the signal source type classifier can include: 4 serially connected fully connected layers. The signal source type classifier is connected to the above-mentioned shallow image feature extraction model, that is, the signal source type classifier takes the output of the shallow image feature extraction model as input. The output size of the last fully connected layer in the 4 serially connected fully connected layers is 1x3, which respectively corresponds to the confidence of the pet living body type, the non-pet living body type and the non-living body type
[0090] Thirdly, in response to the above-mentioned signal source type being the pet living body type, the region of interest is determined according to the above-mentioned signal source position.
[0091] The above-mentioned region of interest includes: at least one superimposed region of interest.
[0092] In practice, first, the signal source position is the three-dimensional coordinates of the signal source in the earth coordinate system (XYZ coordinate system). Therefore, it is necessary to map the signal source position to the coordinates in the image coordinate system (UV coordinate system) through coordinate conversion. Since the main camera 5 and the auxiliary camera 6 included in the video acquisition device are calibrated before use, the signal source position can be mapped to the image coordinate system through coordinate conversion to obtain the coordinate representation of the signal source position in the image coordinate system as the target signal source position. Secondly, the above-mentioned execution subject can construct K regions of interest with the target signal source position as the center. Wherein, K is greater than or equal to 3. Since the K regions of interest are all centered on the target signal source position, the K regions of interest are in a superimposed relationship, that is, there is an intersection between any two regions of interest, so it can be understood as at least one superimposed region of interest.
[0093] Fourthly, according to the above-mentioned at least one superimposed region of interest, the feature of each shallow image feature in the above-mentioned shallow image feature sequence is windowed to generate a windowed shallow image feature, and a windowed shallow image feature sequence is obtained.
[0094] In practice, for the shallow image feature, the shallow image feature can be divided into K layers, each layer corresponding to a superimposed region of interest. Wherein, K is the number of superimposed regions of interest in the at least one superimposed region of interest. The feature values contained in each layer of the superimposed region of interest are unchanged, and the feature values outside the superimposed region of interest in the layer are set to 0. Then the K processed layers are taken as the windowed shallow image feature corresponding to the shallow image feature.
[0095] Fifthly, the above-mentioned windowed shallow image feature sequence is subjected to deep image feature extraction to generate a deep image feature, and a deep image feature sequence is obtained.
[0096] The deep image features can be image features of the monitoring image under a low receptive field. In practice, the execution subject can extract deep image features from the windowed shallow image feature sequence by using a deep image feature extraction model to generate deep image features and obtain a deep image feature sequence.
[0097] As an example, the deep image feature extraction model adopts a feature pyramid network. Specifically, the deep image feature extraction model adopts a down-sampling network composed of five convolutional layers to extract deep image features from the windowed shallow image feature sequence to generate deep image features.
[0098] In the sixth step, the pet type is determined according to the deep image feature sequence.
[0099] In practice, the execution subject can determine the pet type according to the deep image feature sequence by using a pet type classifier. The pet type classifier is a multi-classifier. The shallow image feature extraction model, the signal source type classifier, the deep image feature extraction model, and the pet type classifier are included in a pet recognition model.
[0100] Optionally, after determining the pet type, the pet identity can be associated with the deep image features to generate a labeled training sample. In addition, cases of missed recognition and / or misrecognition within a fixed time period can also be collected as negative (training) samples, which can achieve the purpose of sample optimization. At the same time, by continuously generating training samples through pet images collected under different environments, lighting conditions, and shooting angles, the purpose of continuously enriching the training sample library can be achieved, which can improve the robustness and accuracy of the pet type classifier. In addition, it can also achieve the purpose of a personalized model.
[0101] In step 704, in response to the pet type and the audio description information satisfying the first condition, the pet state description information is determined according to the signal source position and the monitoring audio and video.
[0102] In some embodiments, the execution subject can determine the pet state description information according to the signal source position and the monitoring audio and video in response to the pet type and the audio description information satisfying the first condition, which can include the following steps:
[0103] In the first step, video features are extracted from the monitoring video to generate video features.
[0104] The video feature extraction model shares the shallow image feature extraction model and the deep image feature extraction model included in the pet recognition model. In addition, to avoid repeated extraction of features, the deep image feature sequence can be directly used as the video features.
[0105] Secondly, audio feature extraction is performed on the monitoring audio to generate audio features.
[0106] The audio feature extraction model is an audio feature extraction model in the audio description information generation model. In addition, to avoid repeated extraction of features, the audio features obtained by performing audio feature extraction on the monitoring audio using the audio feature extraction model can be directly used as the audio features obtained in the second step.
[0107] Thirdly, the video features and the audio features are fused to generate fusion features.
[0108] In practice, the execution subject can fuse the video features and the audio features using a fusion model to generate fusion features. Specifically, to fuse the video features and the audio features, the fusion model of the present disclosure uses a cross model attention based on an attention mechanism as the fusion model to fuse the video features and the audio features. Specifically, the fusion model includes a cross model attention A based on an attention mechanism and a feedforward network A, and a cross model attention B based on an attention mechanism and a feedforward network B. The cross model attention A based on an attention mechanism and the feedforward network A are connected in series to extract features from the video features. The cross model attention B based on an attention mechanism and the feedforward network B are connected in series to extract features from the audio features. Then, the extracted video features and the extracted audio features are spliced to generate fusion features.
[0109] Fourthly, the pet state description information is generated according to the fusion features and a pet state description information predictor.
[0110] The pet state description information predictor includes a pet behavior type predictor and a pet health degree predictor. Both the pet behavior type predictor and the pet health degree predictor are multi-classifiers. The pet behavior type predictor is used to classify pet behavior types. For example, pet eating behavior type, pet sleeping behavior type, and pet walking behavior type.
[0111] Step 705, in response to the pet type and the audio description information not satisfying the first condition, determining the pet state description information according to the signal source position and the monitoring video included in the monitoring audio and video.
[0112] In some embodiments, the execution subject described above can determine pet state description information according to the signal source position and the monitoring video included in the monitoring audio and video, in response to the pet type and the audio description information not satisfying the first condition. In practice, since the use of the monitoring audio is not involved at this time, the video features can be directly input as the fusion features into the pet state description information predictor to generate the pet state description information described above. At this time, considering that the feature dimensions of the fusion features containing only the video features and the fusion features containing the video features and the audio features are different. Therefore, before the fusion features are input into the pet state description information predictor, the feature dimension needs to be increased to keep consistent with the feature dimension of the fusion features containing the video features and the audio features.
[0113] In step 706, the pet health degree is accumulated and updated according to the pet behavior type and the pet behavior duration, to obtain an updated pet health degree.
[0114] In some embodiments, the execution subject described above can accumulate and update the pet health degree according to the pet behavior type and the pet behavior duration, to obtain an updated pet health degree.
[0115] In practice, each pet behavior type corresponds to a corresponding base score, the base score is accumulated and biased by the pet behavior duration to obtain a bias score, and the pet health degree is updated to obtain an updated pet health degree.
[0116] In step 707, the health degree abnormality prompt information is sent to the target client in response to the updated pet health degree being located in a first health degree interval.
[0117] In some embodiments, the execution subject described above can send the health degree abnormality prompt information to the target client in response to the updated pet health degree being located in the first health degree interval. The first health degree interval represents a score interval in which the pet is in a health abnormality. In practice, the health degree abnormality prompt information can be "XX, your pet's health degree is XX, and the health degree is low. Please check your pet in time". The target client can be a client corresponding to an account bound to a pet identity corresponding to the pet locator.
[0118] In step 708, the pet behavior abnormality prompt information is sent to the target client in response to the pet behavior type being an abnormal behavior type and the pet behavior duration being greater than an abnormal behavior duration threshold corresponding to the pet behavior type.
[0119] In some embodiments, the execution subject described above can send the pet behavior abnormality prompt information to the target client in response to the pet behavior type being an abnormal behavior type and the pet behavior duration being greater than an abnormal behavior duration threshold corresponding to the pet behavior type.
[0120] In practice, the pet behavior abnormality prompt information can be "Hello XX, your pet is currently behaving abnormally and for a long time, please pay attention to your pet in time".
[0121] Optionally, the method further comprises:
[0122] First, determine whether to allow video uploading.
[0123] In practice, in order to alleviate the local storage pressure and ensure the effectiveness of video backtracking, it is necessary to upload the video. However, since the monitoring video is a monitoring video for a user's residence or other area, in order to avoid user privacy leakage, therefore, before uploading the video, it is necessary to confirm the video uploading behavior. Specifically, a video uploading permission confirmation can be sent to the target client bound to the user. When the user permits video uploading, the monitoring video can be uploaded. When the user does not permit video uploading, only local storage of the monitoring video is allowed.
[0124] Second, in response to the permission, the monitoring video is partially blurred to obtain a partially blurred monitoring video.
[0125] In practice, the execution subject can perform Gaussian blur on the private area in the monitoring video to obtain the partially blurred monitoring video.
[0126] In some optional implementations of some embodiments, the execution subject performs partial video blurring on the monitoring video to obtain a partially blurred monitoring video, comprising:
[0127] First, determine the region confidence of each superimposed region of interest in the at least one superimposed region of interest.
[0128] The region confidence represents the confidence that the corresponding superimposed region of interest contains a pet.
[0129] In practice, when the pet type classifier outputs the pet type, it will obtain the confidence for different pet types, so the pet type classifier can be used to determine the region confidence corresponding to each superimposed region of interest.
[0130] Second, according to the region size of the superimposed region of interest and the region confidence corresponding to the superimposed region of interest, the superimposed region of interest that meets the second condition is selected from the at least one superimposed region of interest as the target region of interest.
[0131] The second condition is that the region size of the superimposed region of interest is relatively large and the corresponding region confidence is relatively large.
[0132] A third sub-step, according to the target region of interest, the shallow image feature sequence is cropped to generate a cropped shallow image feature, and a cropped shallow image feature sequence is obtained.
[0133] In practice, the execution subject can set the feature value of the target region of interest in the shallow image feature to 0 to obtain the cropped shallow image feature.
[0134] A fourth sub-step, according to the cropped shallow image feature sequence and the non-pet living body positioning model, the position of the non-pet living body type object in the monitoring video is determined, and a non-pet living body position is obtained.
[0135] In practice, the non-pet living body positioning model can use Lite-Uet model, which can realize accurate positioning, and the lightweight model can reduce data processing pressure.
[0136] A fifth sub-step, video local blur is performed at the non-pet living body position in the monitoring video to obtain a local blurred monitoring video.
[0137] In practice, the execution subject can perform video local blur at the non-pet living body position in the monitoring video by using a Gaussian blur algorithm to obtain a local blurred monitoring video.
[0138] A third step, video compression is performed on the local blurred monitoring video to generate a compressed video.
[0139] In practice, the execution subject can perform video compression on the local blurred monitoring video by using an H.265 algorithm to generate a compressed video.
[0140] A fourth step, local and cloud dual-end storage is performed on the compressed video.
[0141] In practice, the execution subject can upload the compressed video to a cloud server and store it locally.
[0142] Optionally, data encryption is performed on the transmitted data during data transmission.
[0143] For example, video encryption is performed on the compressed video during uploading to the cloud. For another example, signal encryption is performed on the positioning signal when the pet recognition device sends the positioning signal.
[0144] Through data encryption transmission and video blurring of the person in the monitoring video, the protection of user privacy data is realized, and the trust and satisfaction of users are improved.
[0145] From Figure 7 It can be seen that, compared with Figure 1 the description of some embodiments corresponding to the present disclosure, firstly, the way of determining the signal source position is refined, and the positioning accuracy is improved. Secondly, by combining the pet behavior description information, the pet health degree and abnormal behavior are monitored in real time, and timely warning is given. Thirdly, on the premise of obtaining user upload permission confirmation, in order to alleviate the local storage pressure and ensure the effectiveness based on video backtracking, the user privacy is avoided to be leaked by the way of video local blur.
[0146] Reference will now be made to Figure 8 , which shows a structural schematic diagram of a pet recognition equipment (for example, a computing device) 800 suitable for being used to implement some embodiments of the present disclosure. Figure 8 The pet recognition equipment shown is only an example, and should not bring any limitation to the function and use range of the embodiments of the present disclosure.
[0147] As Figure 8 shown, the pet recognition equipment 800 can include a processing device (for example, a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to the programs stored in a read-only memory 802 or loaded into a random access memory 803 from a storage device 808. In the random access memory 803, various programs and data required for the operation of the pet recognition equipment 800 are also stored. The processing device 801, the read-only memory 802 and the random access memory 803 are connected to each other through a bus 804. An input / output interface 805 is also connected to the bus 804.
[0148] In addition, the pet recognition equipment can also include a pet locator, wherein the pet locator includes at least two ultra-wideband signal transmitters, and the pet locator is fixed to the surface of the pet in a wearable form. A signal collection device, wherein the signal collection device includes at least three ultra-wideband signal receivers, a video collection device and an audio collection device, and the at least three ultra-wideband signal receivers, the video collection device and the audio collection device are oriented to the same collection surface. In practice, the ultra-wideband signal transmitters are used to transmit positioning signals. The video collection device and the audio collection device are used to collect monitoring audio and video. The ultra-wideband signal receivers are used to receive positioning signals.
[0149] In general, the following devices can be connected to the input / output interface 805: input devices 806 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 808 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 809. The communication devices 809 can allow the pet identification apparatus 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 The pet identification apparatus 800 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or fewer devices can alternatively be implemented or present. Figure 8 Each block shown in the flowcharts can represent a device or multiple devices as needed.
[0150] In particular, processes described above with reference to the flowcharts can be implemented as a computer software program according to some embodiments of the present disclosure. For example, some embodiments of the present disclosure include a computer program product including a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In some such embodiments, the computer program can be downloaded and installed from a network through the communication devices 809, or installed from the storage devices 808, or installed from the read-only memory 802. When the computer program is executed by the processing devices 801, the above-described functions defined in the methods of some embodiments of the present disclosure are performed.
[0151] Note that the computer-readable medium in some embodiments of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example and without limitation, be an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a computer-readable program code contained in a data signal communicated in a baseband or as part of a carrier wave. Such a propagated data signal can take any of a variety of forms, including but not limited to electro-magnetic, optical, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF (radio frequency), and the like, or any suitable combination of the foregoing.
[0152] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0153] The computer readable medium can be included in the pet identification equipment or exist separately and not assembled in the pet identification equipment. The computer readable medium carries one or more programs, which, when executed by the pet identification equipment, cause the pet identification equipment to: in response to receiving at least one positioning signal sent by the pet identification device, determine a signal source position according to the at least one positioning signal; collect monitoring audio and video in real time according to the signal source position; determine a pet type according to the signal source position and monitoring video included in the monitoring audio and video; in response to the pet type and audio description information satisfying a first condition, determine pet state description information according to the signal source position and the monitoring audio and video, wherein the audio description information is an audio description of monitoring audio included in the monitoring audio and video, and the pet state description information includes pet behavior description information and pet health description information; and in response to the pet type and the audio description information not satisfying the first condition, determine the pet state description information according to the signal source position and monitoring video included in the monitoring audio and video.
[0154] Computer program code for carrying out operations of some embodiments of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0155] The computer program product of the first aspect can further include one or more computer-readable non-transitory storage media having stored thereon computer-executable instructions that, when executed by one or more computer processors, cause the one or more computer processors to perform operations to implement the method of the first aspect. The computer program product of the first aspect can further include a computer readable medium having stored thereon data that can cause a processor or computer system to implement the method of the first aspect.
[0156] The units described in some embodiments of the present disclosure can be implemented by hardware, software, or a combination thereof. The described units can also be located in a processor,
[0157] The functions described above in the present disclosure can be performed at least in part by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0158] The above description is merely illustrative of the exemplary embodiments of the present disclosure and the technical principles of the application. It should be understood by those skilled in the art that the scope of the application involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the embodiments of the present disclosure (but not limited to) having similar functions.
Claims
1. A pet state recognition method applied to a pet recognition device, comprising: determining a signal source position according to at least one positioning signal in response to receiving the at least one positioning signal sent by the pet recognition device; collecting monitoring audio and video in real time according to the signal source position; determining a pet type according to the signal source position and monitoring video included in the monitoring audio and video, comprising: performing shallow image feature extraction on the monitoring video to obtain a shallow image feature sequence; determining a signal source type according to the shallow image feature sequence, wherein the signal source type comprises a living body type and a non-living body type, and the living body type comprises a pet living body type and a non-pet living body type; in response to the signal source type being the pet living body type, mapping the signal source position in a geodetic coordinate system to an image coordinate system as a target signal source position, and determining a region of interest centered on the target signal source position, wherein the region of interest comprises at least one superimposed region of interest; performing feature windowing on each shallow image feature in the shallow image feature sequence according to the at least one superimposed region of interest to generate a windowed shallow image feature, thereby obtaining a windowed shallow image feature sequence; performing deep image feature extraction on the windowed shallow image feature sequence to generate a deep image feature, thereby obtaining a deep image feature sequence; determining the pet type according to the deep image feature sequence; in response to the pet type and audio description information satisfying a first condition, determining pet state description information according to the signal source position and the monitoring audio and video, wherein the audio description information is an audio description of monitoring audio included in the monitoring audio and video, the pet state description information comprises pet behavior description information and pet health description information, and the first condition is that an average pet volume of a pet represented by the pet type is greater than or equal to a preset pet volume, and an audio quality of the monitoring audio represented by the audio description information is greater than a preset audio quality, the monitoring audio including a sound from a source; in response to the pet type and the audio description information not satisfying the first condition, determining the pet state description information according to the signal source position and monitoring video included in the monitoring audio and video.
2. The method of claim 1, wherein, determining the pet type according to the signal source position and the monitoring video included in the monitoring audio and video, comprising: determining a pet type according to the signal source position, the monitoring video included in the monitoring audio and video, and a pet identity corresponding to the positioning signal, and binding the pet identity with a recognized pet target.
3. The method of claim 1, wherein, determining the signal source position according to the at least one positioning signal, comprising: for each positioning signal in the at least one positioning signal, performing the following processing steps: determining a target information set corresponding to the positioning signal, wherein the target information in the target information set comprises a signal time difference, a first coordinate, and a second coordinate, and the signal time difference represents a time difference between the first coordinate corresponding signal receiver and the second coordinate corresponding signal receiver receiving the positioning signal, According to the target information set, a candidate signal source position is determined; According to the obtained candidate signal source position set, signal source position correction is performed to generate the signal source position.
4. The method of claim 1, wherein, The pet behavior description information includes a pet behavior type and a pet behavior duration, and the pet health description information includes a pet health degree. The method further includes: According to the pet behavior type and the pet behavior duration, the pet health degree is accumulated and updated to obtain an updated pet health degree; In response to the updated pet health degree being located in a first health degree interval, health degree abnormality prompt information is sent to a target client; In response to the pet behavior type being an abnormal behavior type and the pet behavior duration being greater than an abnormal behavior duration threshold corresponding to the pet behavior type, pet behavior abnormality prompt information is sent to the target client.
5. The method of claim 1, wherein, The audio description information includes audio source description information and audio quality description information, and the audio source description information represents a sound source included in the monitoring audio. The audio description information is generated by the following steps: Audio feature extraction is performed on the monitoring audio to generate audio features; The audio source description information is generated by an audio source recognition model and the audio features; The audio quality description information is generated according to an audio quality scoring model, the audio features, and the audio source description information.
6. The method of claim 1, wherein, The pet state description information is determined according to the signal source position and the monitoring audio and video, including: Video feature extraction is performed on the monitoring video to generate video features; Audio feature extraction is performed on the monitoring audio to generate audio features; Feature fusion is performed on the video features and the audio features to generate fused features; The pet state description information is generated according to the fused features and a pet state description information predictor, wherein the pet state description information predictor includes a pet behavior type predictor and a pet health degree predictor.
7. The method of claim 1, wherein, The method further includes: Determining whether to allow video uploading; In response to allowing, performing video local blurring on the monitoring video to obtain a locally blurred monitoring video; Performing video compression on the locally blurred monitoring video to generate a compressed video; Respectively performing local-end and cloud-end dual-end storage on the compressed video.
8. The method of claim 7, wherein, The locally blurred monitoring video is obtained by performing video local blurring on the monitoring video, including: Determining a region confidence of each superimposed region of interest in the at least one superimposed region of interest, wherein the region confidence represents a confidence that the pet is included in the corresponding superimposed region of interest; According to the region size of the superimposed region of interest and the region confidence corresponding to the superimposed region of interest, a superimposed region of interest that satisfies a second condition is selected from the at least one superimposed region of interest as a target region of interest; According to the target region of interest, each shallow image feature in the shallow image feature sequence is inversely selected and cropped to generate a cropped shallow image feature, and a cropped shallow image feature sequence is obtained. According to the cropped shallow image feature sequence and the non-pet living body positioning model, a position of an object of a non-pet living body type in the monitoring video is determined, and a non-pet living body position is obtained. Video local blur is performed on the non-pet living body position in the monitoring video, and a local blurred monitoring video is obtained.
9. A pet identification apparatus, comprising: a pet locator, wherein the pet locator comprises at least two ultra-wideband signal transmitters, and the pet locator is fixed to a pet body surface in a wearable form; a signal collection device, wherein the signal collection device comprises at least three ultra-wideband signal receivers, a video collection device, and an audio collection device, and the at least three ultra-wideband signal receivers, the video collection device, and the audio collection device are directed to the same collection surface; one or more processors; a storage device having one or more programs stored thereon; when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 8.
10. A computer readable medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Pet health monitoring method and device, computer equipment and readable storage medium
CN114694842A
Method and device for determining animal species
CN115035450A
Pet activity range generation method and device and related equipment
CN116071232A
Intelligent sensing necklace for pets and control method of intelligent sensing necklace
CN116711654A