Image recognition method and device, communication device and readable storage medium

By extracting and recognizing features from multiple frames of images, feature vectors of motion information, temporal information, and spatial information are obtained, solving the problem of misjudgment by infrared sensors under slight movement and achieving higher recognition accuracy.

CN116152906BActive Publication Date: 2026-04-24CHINA MOBILE COMM LTD RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA MOBILE COMM LTD RES INST
Filing Date
2021-11-18
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In existing lighting control solutions, when using infrared sensors to identify the movement of people or objects, misjudgments are prone to occur when the person is stationary or makes only slight movements, resulting in low recognition accuracy.

Method used

By acquiring image sequences of the target location, multi-frame image feature extraction is performed. A first feature vector representing movement information and a second feature vector representing temporal and spatial information are obtained. Object recognition is performed by combining the first and second feature vectors to avoid misjudgment in the case of slight movement.

Benefits of technology

It improves the accuracy of recognition, effectively identifies subtle movements, and reduces misjudgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116152906B_ABST
    Figure CN116152906B_ABST
Patent Text Reader

Abstract

The application provides an image recognition method and device, communication equipment and a readable storage medium. The method comprises: acquiring an image sequence of a target position, the image sequence comprising a plurality of first images; performing feature extraction on the plurality of first images to obtain a first feature vector, the first feature vector being used to represent movement information in the plurality of first images; performing feature extraction on the plurality of first images to obtain a second feature vector, the second feature vector being used to represent time domain information and space domain information of the plurality of first images; and performing object recognition on the plurality of first images based on the first feature vector and the second feature vector. The application can improve recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image recognition method, apparatus, communication device, and readable storage medium. Background Technology

[0002] Currently, lighting control solutions typically identify the presence of people or moving objects by installing sensors that detect human presence, such as infrared sensors, to achieve remote lighting control. Infrared sensors trigger signals based on the difference between human body temperature and environmental changes to identify whether someone is present. However, when a person is stationary or makes only slight movements, it is easy to misjudge that no one is present, resulting in a low accuracy rate. Summary of the Invention

[0003] This application provides an image recognition method, apparatus, communication device, and readable storage medium to solve the problem of low recognition accuracy.

[0004] In a first aspect, embodiments of this application provide an image recognition method, including:

[0005] Acquire an image sequence of the target location, the image sequence comprising multiple frames of a first image;

[0006] Feature extraction is performed on the multiple frames of the first image to obtain a first feature vector, which is used to represent the motion information in the multiple frames of the first image.

[0007] Feature extraction is performed on the multiple frames of the first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multiple frames of the first image;

[0008] Based on the first feature vector and the second feature vector, object recognition is performed on the multiple frames of the first image.

[0009] Secondly, embodiments of this application also provide an image recognition device, comprising:

[0010] The first acquisition module is used to acquire an image sequence of the target location, the image sequence including multiple frames of the first image;

[0011] A first extraction module is used to extract features from the multi-frame first image to obtain a first feature vector, the first feature vector being used to represent motion information in the multi-frame first image;

[0012] The second extraction module is used to extract features from the multiple frames of the first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multiple frames of the first image.

[0013] The recognition module is used to perform object recognition on the multi-frame first image based on the first feature vector and the second feature vector.

[0014] Thirdly, embodiments of this application also provide a communication device, including: a transceiver, a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the method described in the first aspect of embodiments of this application.

[0015] Fourthly, embodiments of this application also provide a readable storage medium storing a program that, when executed by a processor, implements the steps in the method described in the first aspect of embodiments of this application.

[0016] In this embodiment, feature extraction is performed on the multiple first frames to obtain a first feature vector, which represents the motion information in the multiple first frames; feature extraction is also performed on the multiple first frames to obtain a second feature vector, which represents the temporal and spatial information of the multiple first frames; based on the first and second feature vectors, object recognition is performed on the multiple first frames. The presence of minute movements can be determined by the motion information and the temporal and spatial information of the multiple first frames, thereby avoiding misjudgment in the case of minute movements and improving the accuracy of recognition. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of this application, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic flowchart of an image recognition method provided in an embodiment of this application;

[0019] Figure 2 This is a schematic diagram of pixels in a multi-frame image provided in an embodiment of this application;

[0020] Figure 3 This is a schematic diagram of a lighting automation control system provided in an embodiment of this application;

[0021] Figure 4 This is a schematic diagram of pixel block extraction provided in an embodiment of this application;

[0022] Figure 5 This is a schematic diagram illustrating the assignment of image pixel values ​​according to an embodiment of this application;

[0023] Figure 6 This is a schematic diagram illustrating another method of assigning image pixel values ​​according to an embodiment of this application;

[0024] Figure 7 This is a schematic diagram of an image feature extraction process provided in an embodiment of this application;

[0025] Figure 8 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0026] Figure 9 This is a schematic diagram of the structure of a communication device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in the embodiments of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices. Additionally, the use of "and / or" in this application indicates at least one of the connected objects, such as A and / or B and / or C, representing seven possibilities: including A alone, B alone, C alone, and the presence of both A and B, both B and C, both A and C, and the presence of A, B, and C.

[0029] Please see Figure 1 , Figure 1 This is a flowchart illustrating an image recognition method provided in an embodiment of this application, as shown below. Figure 1 As shown, it includes the following steps:

[0030] Step 101: Obtain the image sequence of the target location, the image sequence including multiple frames of the first image.

[0031] The target location can be selected in a pre-defined area, for example, by selecting the size of the target location in pixels, thereby obtaining multiple frames of images of the same size.

[0032] It can be understood that the above-mentioned first multi-frame image is a multi-frame image acquired according to a preset time interval, for example: acquiring multiple frames of images at a frequency of one frame every 10 seconds, that is, the time interval between adjacent frames is 10 seconds, thereby capturing multiple frames of images of the target location at different times.

[0033] Optionally, before obtaining the image sequence of the target location in step 101, the method may further include the following steps:

[0034] Multiple parallel threads are created to acquire multiple frames of third images.

[0035] Obtain the region of interest (ROI) of the multi-frame third image, wherein the target location is located within the ROI;

[0036] The step 101 of obtaining the image sequence of the target location may specifically include:

[0037] The image sequence of the target location is obtained based on the multi-frame third image.

[0038] The Region of Interest (ROI) can be predetermined. Specifically, the image acquisition device is usually located at a fixed position, and the area covered by the images that the device can acquire is also determined. Therefore, the area within the coverage area can be pre-marked as the ROI to improve the efficiency and accuracy of image recognition. For example, in a conference room, the seating area next to the conference table can be marked as the ROI. By selectively recognizing multiple frames of images acquired from the ROI, the amount of data and accuracy of image recognition can be reduced.

[0039] The aforementioned region of interest may include at least one of the aforementioned target regions. For example, in a conference room, each location may be marked as a region of interest. The target location in a region of interest may be the entire region of interest, or a region of interest may include multiple target locations divided according to pixel size.

[0040] The number of the aforementioned multi-frame third images may differ from the number of the aforementioned multi-frame first images. For example, when the image recognition is performed in real time, the acquisition time of the most recent frame of the aforementioned multi-frame third images is real time, while the other images are historical images of adjacent frames. The aforementioned multi-frame third images can be acquired by the aforementioned multiple parallel threads, thereby obtaining the image sequence of the aforementioned target location and realizing real-time image recognition. When the image recognition is performed according to a preset time interval, the aforementioned multiple parallel threads can acquire and store the acquired multi-frame third images. When the execution time of image recognition arrives, the image sequence of the aforementioned target location is obtained based on the stored multi-frame third images. The number of the aforementioned multi-frame third images is greater than the number of multiple first images required for image recognition, thereby realizing the current image recognition. When the execution time of the next image recognition arrives, the above steps are repeated. It can be understood that the aforementioned image sequence is acquired starting from the most recent frame.

[0041] In this embodiment, multiple parallel threads are created to acquire multiple frames of third images, the regions of interest of the multiple frames of third images are acquired, the target location is located within the regions of interest, and the image sequence of the target location is acquired based on the multiple frames of third images. This allows for recognition of images within the regions of interest, reducing the amount of image data to be recognized and improving the efficiency of image recognition.

[0042] Step 102: Extract features from the multi-frame first image to obtain a first feature vector, which is used to represent the motion information in the multi-frame first image.

[0043] The aforementioned movement information may be movement information of objects or people moving in the target location, such as movement direction information from the lower left corner of the image to the lower right corner of the image. If there are no objects or people moving in the target location, the aforementioned movement information may also include information indicating no movement direction.

[0044] Optionally, step 102, which involves extracting features from the multiple frames of the first image to obtain a first feature vector, includes:

[0045] Using multiple pre-set second masks, the pixel intensity of each pixel in the multiple frames of the first image is extracted in the multiple preset directions;

[0046] Obtain the third pixel and the fourth pixel in the first image of the multiple frames, wherein the third pixel is the pixel with the largest pixel intensity among the pixel intensities of the multiple preset directions, and the fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities of the multiple preset directions.

[0047] The first feature vector is constructed based on the third pixel and the fourth pixel.

[0048] It is understood that the aforementioned multiple second masks can correspond one-to-one with the aforementioned multiple preset directions. For example, the pixels of the aforementioned multiple frames of the first image can constitute a three-dimensional pixel block. Eight of the aforementioned second masks can be set for the eight vertices of this three-dimensional pixel block to extract the pixel intensity of each pixel in the aforementioned multiple frames of the first image after processing by the aforementioned eight second masks. The aforementioned third pixel is the pixel with the largest pixel intensity among the pixel intensities of the multiple preset directions, and the aforementioned fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities of the multiple preset directions. That is, the aforementioned third pixel is the starting position with the largest gradient change, and the aforementioned fourth pixel is the ending position with the largest gradient change. Thus, the object movement direction information in the aforementioned multiple frames of the first image can be fused based on the positions of the aforementioned third pixel and the aforementioned fourth pixel. The object movement direction information can be represented by the aforementioned first feature vector. Specifically, by recognizing the object movement direction in the aforementioned multiple frames of the first image, the recognition of minute movements can be achieved, avoiding image recognition errors under minute movements.

[0049] Specifically, the process of using multiple second masks to extract the pixel intensity of each pixel in the multiple frames of the first image in multiple preset directions can be understood as multiple convolution templates convolving with the three-dimensional matrix formed by each pixel in the multiple frames of the first image, thereby assigning a value to each pixel in the multiple frames of the first image, extracting the pixel intensity of each pixel in the multiple frames of the first image in the corresponding preset direction, and increasing robustness to noise.

[0050] In this embodiment, by using multiple pre-set second masks, the pixel intensity of each pixel in the multiple frames of the first image is extracted in multiple preset directions; a third pixel and a fourth pixel in the multiple frames of the first image are obtained, wherein the third pixel is the pixel with the highest pixel intensity in the multiple preset directions, and the fourth pixel is the pixel with the lowest pixel intensity in the multiple preset directions; a first feature vector is constructed based on the third pixel and the fourth pixel; the movement direction of a minute action can be obtained by extracting the pixel intensity in the multiple preset directions.

[0051] Step 103: Extract features from the multi-frame first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multi-frame first image.

[0052] The aforementioned temporal information can be understood as the temporal relationship information of the same pixel in the first image across different frames, and the aforementioned spatial information can be understood as the spatial relationship information between pixels in a frame of the first image. For example, taking the image sequence as including three frames of the first image, and the acquisition times of the three frames of the first image as t, t+1 and t+2 respectively, temporal information can be obtained for the same pixel at t, t+1 and t+2 respectively, and spatial information can be obtained for different pixels in one of the frames of the image.

[0053] Optionally, the feature extraction of the multiple frames of the first image in step 103 to obtain the second feature vector may specifically include the following steps:

[0054] Obtain the center pixel of the first image in multiple frames, and a plurality of third pixels surrounding the center pixel;

[0055] The second feature vector is constructed based on the center pixel and the plurality of third pixels.

[0056] Specifically, such as Figure 2 As shown, these are three frames taken at the same location, representing frames n-1, n, and n+1 respectively. The three frames are represented by pixels p. j Center pixel and For the center pixel p j The pixels, as can be understood, the above and as well as and It contains spatial pixels, representing spatial information. and Including temporal information from previous and next frames, feature vectors representing the temporal and spatial information can be obtained by extracting features from the three frames based on each pixel.

[0057] In this embodiment, by obtaining the center pixel of the multi-frame first image and a plurality of third pixels surrounding the center pixel, and constructing the second feature vector based on the center pixel and the plurality of third pixels, the extraction of the second feature vector can be achieved.

[0058] Optionally, constructing the second feature vector based on the center pixel and the plurality of third pixels may specifically include:

[0059] Obtain multiple differences between the center pixel and the plurality of third pixels;

[0060] The second feature vector is constructed using a step function and the plurality of differences.

[0061] The aforementioned multiple differences may include the difference between the grayscale value of the center pixel and the average grayscale value of the multiple third pixels, and the grayscale value differences between the third pixels that constitute temporal or spatial information. Specifically, for example... Figure 2 As shown, through and The difference in grayscale values ​​can represent the temporal information of an image, through... and The difference in grayscale values, and and The difference in gray values ​​can represent the spatial information of an image, and multiple gray value differences can be represented using a step function to reduce the feature dimension.

[0062] In this embodiment, by obtaining multiple differences between the center pixel and the multiple third pixels, the connection between the center pixel and the multiple third pixels surrounding the center pixel is enhanced. The second feature vector is constructed using a step function and the multiple differences, which can reduce the dimension of the second feature vector, reduce the amount of data processed for image recognition, and improve the efficiency of image recognition.

[0063] Step 104: Based on the first feature vector and the second feature vector, perform object recognition on the multi-frame first image.

[0064] The object recognition mentioned above can include the recognition of moving objects at the target location. For example, a person's slight movement or slight action can be understood as the movement of an object.

[0065] It is understood that, based on the first feature vector and the second feature vector, object recognition can be performed on the first multi-frame images. This can be done by identifying whether there is object movement in the first multi-frame images based on the movement information, temporal information, and spatial information in the first multi-frame images.

[0066] The object recognition process for the aforementioned multiple frames of the first image can be achieved through a binary classification solution using a classifier. For example, the feature vectors to be recognized, such as the first and second feature vectors mentioned above, can be concatenated and then input into an SVM (Support Vector Machine) classifier to obtain the classification result of the image recognition, i.e., whether it is identified as having a person or not. Alternatively, other classifiers such as decision trees can be used to perform object recognition on the aforementioned multiple frames of the first image to obtain the classification result.

[0067] In addition, after performing object recognition on the aforementioned multiple frames of the first image, the recognition results can be output. If the image recognition method is applied to a conference room, the lighting in the conference room can be controlled based on the recognition results.

[0068] In this embodiment, feature extraction is performed on the multiple first frames to obtain a first feature vector, which represents the motion information in the multiple first frames; feature extraction is also performed on the multiple first frames to obtain a second feature vector, which represents the temporal and spatial information of the multiple first frames; based on the first and second feature vectors, object recognition is performed on the multiple first frames. The presence of minute movements can be determined by the motion information and the temporal and spatial information of the multiple first frames, thereby avoiding misjudgment in the case of minute movements and improving the accuracy of recognition.

[0069] Optionally, before performing object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector in step 104, the method may further include the following steps:

[0070] Delete the edge pixels of the multi-frame first image to generate multi-frame second image;

[0071] Feature extraction is performed on the multiple frames of the second image to obtain a third feature vector, which is used to represent the motion information in the multiple frames of the second image;

[0072] Step 104, which involves performing object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector, may specifically include:

[0073] The first feature vector, the second feature vector, and the third feature vector are concatenated to obtain the target feature vector;

[0074] Based on the target feature vector, object recognition is performed on the multi-frame first image.

[0075] It is understandable that since the boundary pixels of the middle frame in the spatial domain do not participate in the extraction of the above pixel intensity, and are located in the middle position in the temporal domain, they are not the main pixels affecting the change of directional gradient. Therefore, by deleting the above edge pixels, the dimension of the above third feature vector can be reduced by performing feature extraction on the generated multi-frame second image. Furthermore, by concatenating the first feature vector, the second feature vector, and the third feature vector, a target feature vector can be obtained. Based on the target feature vector, object recognition is performed on the multi-frame first image, thereby enhancing the discriminative power of the target feature vector.

[0076] In this embodiment, edge pixels of the multi-frame first image are deleted to generate multi-frame second images; features are extracted from the multi-frame second images to obtain a third feature vector, which represents the motion information in the multi-frame second images; the first feature vector, the second feature vector, and the third feature vector are concatenated to obtain a target feature vector; based on the target feature vector, object recognition is performed on the multi-frame first images, which has richer information during object recognition, and the third feature vector has a lower dimension, which can improve the discriminative power of the target feature vector.

[0077] Optionally, the step of extracting features from the multiple frames of the second image to obtain a third feature vector may specifically include:

[0078] Using multiple pre-set first masks, the pixel intensity of each pixel in the multiple frames of the second image is extracted in multiple preset directions;

[0079] Obtain the first pixel and the second pixel in the multiple frames of the second image, wherein the first pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the second pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions.

[0080] The third feature vector is constructed based on the first pixel and the second pixel.

[0081] It is understood that the aforementioned multiple first masks can correspond one-to-one with the aforementioned multiple preset directions. For example, the pixels of the aforementioned multiple frames of first images can constitute a three-dimensional pixel block. For the eight vertices of this three-dimensional pixel block, eight of the aforementioned first masks can be set respectively to extract the pixel intensity of each pixel in the aforementioned multiple frames of second images after processing by the aforementioned eight first masks. The aforementioned first pixel is the pixel with the largest pixel intensity among the pixel intensities of the multiple preset directions, and the aforementioned second pixel is the pixel with the smallest pixel intensity among the pixel intensities of the multiple preset directions. That is, the aforementioned first pixel is the starting position with the largest gradient change, and the aforementioned second pixel is the ending position with the largest gradient change. Thus, the object movement direction information in the aforementioned multiple frames of second images can be fused based on the positions of the aforementioned first pixel and the aforementioned second pixel. The object movement direction information can be represented by the aforementioned third feature vector. Specifically, by recognizing the object movement direction in the aforementioned multiple frames of second images, the recognition of minute movements can be achieved, avoiding image recognition errors under minute movements.

[0082] In this embodiment, by using multiple pre-set first masks, the pixel intensity of each pixel in the multiple frames of the first image is extracted in multiple preset directions; a first pixel and a second pixel in the multiple frames of the first image are obtained, wherein the first pixel is the pixel with the largest pixel intensity in the multiple preset directions, and the second pixel is the pixel with the smallest pixel intensity in the multiple preset directions; the third feature vector is constructed based on the first pixel and the second pixel; the movement direction of a minute action can be obtained by extracting the pixel intensity in the multiple preset directions.

[0083] Optionally, before performing feature extraction on the multiple frames of the first image in step 102, the method further includes the following steps:

[0084] Obtain the grayscale difference between the multiple frames of the first image;

[0085] The feature extraction of the multiple frames of the first image in step 102 may specifically include:

[0086] When the grayscale difference is less than a preset threshold, feature extraction is performed on the first multi-frame image.

[0087] The grayscale difference mentioned above can be obtained using the Inter Frame Difference Method (IFDM). This involves performing a difference operation on two or three temporally consecutive frames of the first multi-frame image, subtracting the corresponding pixels from each frame, and determining the absolute value of the grayscale difference. When the absolute value exceeds the preset threshold, the target can be identified as a moving target, thus achieving the object movement detection function.

[0088] When the absolute value of the grayscale value is less than the preset threshold, it can be understood that there is no large movement at the target position. Therefore, by extracting features from the multiple frames of the first image, the first feature vector and the second feature vector are obtained respectively. The first feature vector and the second feature vector are then used to perform object recognition on the multiple frames of the first image to identify whether there is any movement of minute motion, thereby improving the accuracy of recognition.

[0089] In this embodiment, before performing feature extraction on the multi-frame first images, the grayscale difference between the multi-frame first images is obtained. If the grayscale difference is less than a preset threshold, feature extraction is performed on the multi-frame first images. That is, when there is a large movement at the target location, the judgment can be made directly based on the grayscale difference, which has high robustness. When there is a small movement at the target location, object recognition is further performed based on the features extracted from the multi-frame first images.

[0090] The various optional implementation methods described in the embodiments of this application can be combined with each other or implemented individually without conflict. The embodiments of this application do not limit this.

[0091] For ease of understanding, the specific implementation method is as follows:

[0092] This application provides a lighting automation control system, including a pyroelectric infrared sensor, a webcam, and an embedded image computing device. For example... Figure 3 As shown, the lighting automation control system can reduce the false recognition rate of pyroelectric infrared sensors in identifying human bodies by linking data from pyroelectric infrared sensors and network cameras. It can also work with management platforms, smart gateways, and smart switches to further realize automated lighting control.

[0093] The pyroelectric infrared sensor is an infrared radiation detection sensor made of crystal materials with pyroelectric effect, which can effectively detect moving human bodies within the field of view. Its working principle is that when a human body moves in a certain direction and at a certain distance, a voltage output signal is generated, causing a step input to the corresponding pyroelectric element in the sensing area. This, in turn, generates a pyroelectric voltage through subsequent circuitry, realizing the output of a human movement detection signal. It is suitable for detecting people moving at close range. The camera is a front-end network camera for image acquisition. Network cameras are a new generation product combining traditional cameras with network video technology. In addition to possessing all the image capture functions of a typical traditional camera, it also has a built-in digital compression controller and a WEB (World Wide Web) based operating system, enabling video data to be compressed and encrypted before being transmitted to the end user via a local area network, the Internet, or a wireless network. Remote users can access the network camera using a standard web browser on a PC (Personal Computer) based on its IP (Internet Protocol) address. They can monitor the target site in real time, edit and store image data, and control the camera's pan / tilt / zoom (PTZ) for comprehensive monitoring. Multiple image data streams are stored locally on the NVR (Network Video Recorder). The NVR's primary function is to receive, store, and manage digital video streams transmitted from IPC (Internet Protocol Camera) devices over the network, leveraging the distributed architecture advantages of networking. Simply put, an NVR allows simultaneous viewing, browsing, playback, management, and storage of multiple network camera feeds, eliminating the constraints of traditional computer hardware and reducing the hassle of software installation. The embedded image computing device is a microcomputer integrating the image presence detection algorithm model, communication module, and graphics card support described in this proposal. It has a USB (Universal Serial Bus) interface and can be plugged into the NVR for image extraction and computation.

[0094] Specifically, for scenarios with large dynamic range, inter-frame difference is used to identify moving objects; for scenarios with minor motion, based on the data features of images in fixed scenarios, the region of interest is first divided, and features are extracted from 3D features that fuse temporal and spatial information. Finally, a classifier is used to solve a binary classification problem to obtain a classification result of whether someone is present or absent. This application can reduce the possibility of misjudging people in lighting control schemes in scenarios with minor motion, while also considering the possible states of the scene globally, achieving a one-stop convenient solution based on image data. A storage and computing device is also proposed, integrating model computing and communication capabilities. It enables local image computing through plug-and-play NVR, keeping data within the park, and converts image signals into wirelessly encoded signals for transmission to a local gateway to achieve sensor-linked lighting control. Furthermore, it can be transmitted to a platform via the Internet to achieve value-added functions such as remote lighting control and conference room usage statistics. The software part of this application mainly consists of the algorithms mentioned in the following pseudocode:

[0095] Input: pyroelectric signal and image;

[0096] Output: Lighting status (on / off);

[0097]

[0098] This application also provides a personnel identification method applied to a lighting automation control system, which may specifically include the following steps:

[0099] Step 1: When personnel enter the conference room, the pyroelectric infrared sensor located at the door will detect the movement of living beings based on the principle of infrared reflection, and transmit the wireless signal to the gateway, or further to the management platform.

[0100] Step 2: The gateway responds to the linkage logic, turns on the lighting automation control system, starts the network camera, and turns off the pyroelectric infrared sensor.

[0101] Step 3: The webcam continuously captures images at a rate of one frame every 10 seconds.

[0102] Step 4: The network camera uploads the images to the NVR for storage, and further creates parallel threads in the storage medium of the embedded image computing device to read video frames. The embedded image computing device stores the video frames in a queue and retrieves them as needed during detection.

[0103] Step 5: After the embedded image computing device reads in three images, it compares the image pixels using the inter-frame difference method. When a pixel change is detected and exceeds a set threshold, the illumination state remains on; when no pixel change is detected and exceeds the set threshold, proceed to step 6. The inter-frame difference method specifically includes: when there is a moving object in the video, there will be a difference in grayscale between adjacent frames (or three adjacent frames). The absolute value of the grayscale difference between two frames is calculated. A stationary object will appear as 0 in the difference image, while a moving object, especially its outline, will appear non-zero due to grayscale changes. When the absolute value exceeds a certain threshold, it can be identified as a moving target, thus achieving target detection. The inter-frame difference method can be applied to scenes with large dynamic ranges. The algorithm is simple to implement and has low program design complexity; it is not very sensitive to changes in lighting and other scene variations, can adapt to various dynamic environments, and has strong robustness.

[0104] Step 6: Take the three most recent captured frames and perform feature extraction and classification on each image to determine if someone is present at that moment. Specifically, this may include the following process:

[0105] First, since the key to micro-motion recognition lies in obtaining its motion direction information, 5×5 pixels are taken from the same position in three consecutive frames to form a 5×5×3 consecutive frame for motion direction information encoding. For example... Figure 4 and Figure 5 As shown, eight new directional angle masks are designed to find motion information, and the directional angle intensity value Q is calculated based on the eight directional masks. ji (p j ), by pixel p j The pixel values ​​of the central cube form a 3D matrix M, which is then convolved with 8 directional angular masks to obtain:

[0106] Q ji (p j )=M*m j ;

[0107] Where j represents the j-th feature pixel block extracted in the entire image frame, i = 1, 2, ..., 8, * is the convolution symbol, p j m is the center pixel of the cube. j This represents the angular mask for each direction.

[0108] Figure 5 Examples of assignment in two directions are given. Black pixels are assigned a value of 6, white pixels are assigned a value of -1, and stripe pixels, which do not participate in direction angle encoding, are assigned a value of 0. Figure 5Taking the left image as an example, in the first frame, the center pixel (striped area) is assigned a value of 0, the five pixels in the lower left corner (black area) are assigned a value of 6, and the remaining pixels (white area) are assigned a value of -1; in the second frame, the center pixel (striped area) is assigned a value of 0, the middle pixels (striped areas) of the four sides of the image are assigned a value of 0, the lower left pixel (black area) is assigned a value of 6, and the remaining pixels (white area) are assigned a value of -1; in the third frame, the center pixel (striped area) is assigned a value of 0, the middle pixel (striped area) of the two adjacent sides in the upper right corner of the image is assigned a value of 0, and the remaining pixels (white area) are assigned a value of -1, thereby extracting the pixel intensity of each pixel in that direction angle.

[0109] By multiplying the pixel value by the color value assigned to the corresponding position, the directional intensity value can be calculated. The directional intensity value represents the pixel intensity in different directional angle regions, and encoding is performed based on the intensity value changes at the location of the directional angle.

[0110] The positions corresponding to the maximum and minimum intensity values ​​are encoded, that is, the starting and ending positions of the largest gradient changes are encoded, for use in fusing motion direction information. The encoding formula is as follows:

[0111] D j (p j )=Q max (p j )+Q min (p j );

[0112] Among them, Q max (p j ) and Q min (p j ) represent the locations of the maximum and minimum strength values ​​in each of the multiple directions, respectively.

[0113] Secondly, spatial redundancy vectors are removed. Since the boundary pixels of the middle frame in the spatial domain do not participate in the calculation of the directional intensity value and are located in the middle position in the temporal domain, they are not considered as the main pixels affecting the directional gradient change. Next, we consider the spatiotemporal combination of pixels after removing spatial redundancy vectors. Figure 6 Examples of assignments in two directions are given. Black pixels are assigned a value of 4, white pixels are assigned a value of -1, and striped pixels are not involved in directional angle encoding and are assigned a value of 0. Multiplying each pixel value by the corresponding color assignment yields a directional intensity value, which represents the pixel intensity in different directional angle regions. Encoding is performed based on the intensity value changes at the location of the directional angle.

[0114] Furthermore, such as Figure 4 and Figure 6As shown, the angular intensity value Q′ of the new sub-pixel block in each direction is calculated based on the mask in 8 directions. ji (p j ), by pixel p j The pixel values ​​of the central cube form a 3D matrix M, which is then convolved with 8 directional angular masks to obtain:

[0115] Q i (p j )=M*m j ;

[0116] Where j represents the j-th feature pixel block extracted in the entire image frame, i = 1, 2, ..., 8, * is the convolution symbol, p j m is the center pixel of the cube. j This represents the angular mask for each direction.

[0117] The positions corresponding to the maximum and minimum intensity values ​​are encoded, that is, the starting and ending positions of the largest gradient changes are encoded, for use in fusing motion direction information. The encoding formula is as follows:

[0118] D′ j (p j )=Q′ max (p j )+Q′ min (p j );

[0119] Among them, Q′ max (p j ) and Q′ min (p j ) represent the locations of the maximum and minimum strength values ​​in each direction, respectively.

[0120] Finally, for the pixel p surrounding the center point j The encoding consists of 6 pixels. These 6 pixels form a simplified 3D orthogonal plane, where, except for the center pixel, the remaining pixels represent the centers of the temporal and spatial domains, respectively. This pixel encoding is highly representative and can make the information connection between 3D pixels closer.

[0121] like Figure 2 and Figure 6 As shown, these pixels are divided into three groups, namely and and It contains spatial pixels, representing spatial information. and It contains temporal information from consecutive frames in the image sequence. To reduce feature dimensionality while enhancing the connection between the center pixel and its surroundings, each group of pixels is interpolated using the following formula:

[0122]

[0123] In the above formula For the corresponding pixel point The grayscale value, where σ is the step function, is calculated using the following formula:

[0124]

[0125] Finally, we will obtain D. j (p j ), D′ j (p j ), D″ j (p j The features are concatenated and used as the total features of the pixel block composed of three frames of images. Figure 7 The extraction process for a 3D feature extraction method that integrates temporal and spatial information is described. A 5×5 pixel block is taken from the same location at different times in three adjacent frames. The edge intensity values ​​at eight angular directions are calculated, and the maximum and minimum positions are used to obtain the D-value. j (p j Then, the boundary strength value D′ after removing redundant pixels is further calculated. j (p j Finally, D″ is calculated in the simplified three-dimensional orthogonal plane region. j (p j Then, concatenating these three features yields 3D features that fuse temporal and spatial information. If only one frame of an image is considered during feature extraction, or if pixels from multiple frames are simply superimposed, the computational complexity can become excessively high. Furthermore, simple encoding methods are easily affected by noise, resulting in insufficient discriminative power of the features. In this application, the 3D features that fuse temporal and spatial information have lower dimensionality and employ eight convolutional templates to increase robustness to noise, while also possessing richer information. Therefore, the features obtained by the feature extraction method proposed in this application have stronger discriminative power.

[0126] Feature classification: The SVM classifier is used to solve the problem and obtain the classification results for personnel identification.

[0127] Step 7: The detection results of the three images are ORed, and the algorithm processing result is converted into a digital analog signal and transmitted to the gateway for further linkage with other devices. Judgment result: When someone is detected, the lighting is turned on, and the process returns to step 3; when the judgment result is no one, wait 30 seconds, then turn off the lighting and camera, and proceed to step 8.

[0128] Step 8: Turn on the infrared sensor.

[0129] In this embodiment, personnel identification is achieved through a hierarchical algorithm. Specifically, for scenes with large dynamic range, the frame difference method is used to identify moving objects; for scenes with minor movement, the region of interest is divided based on the image data features of the fixed scene, and features are extracted using a 3D feature extraction method based on the fusion of temporal and spatial information proposed in this application. Finally, a classifier is used to solve a binary classification problem. The feature extraction method provided in this application is suitable for computing small sample datasets in fixed scenes, improving algorithm speed while facilitating rapid training and deployment of the algorithm model. Furthermore, the embedded image computing device integrating the method of this application features plug-and-play compatibility with NVRs via a USB interface, enabling the reuse of existing network cameras and further reducing installation and development costs.

[0130] Furthermore, the 3D features obtained by fusing temporal and spatial information in this application have a lower dimensionality, and the use of 8 convolutional templates increases robustness to noise while possessing richer information. The feature extraction method proposed in this application yields 144-dimensional features from three pixel blocks, while directly stacking the three pixel blocks results in 768-dimensional features. Therefore, the feature extraction method proposed in this application yields features with a lower dimensionality, which facilitates faster computation, and the features obtained by fusing temporal and spatial information have stronger discriminative power.

[0131] When applied to remote conference room management scenarios, this application uses data linkage between pyroelectric infrared sensors and cameras to determine the presence of personnel and further control lighting. Simultaneously, a gateway converts the image-based determination of the current conference room usage status (on / off) into a binary signal and transmits it to the management platform for remote management, such as conference room usage management. This application reduces the problem of false personnel detection by infrared sensors in scenarios with slight movements, while comprehensively considering the possible states of the conference room, thus achieving remote conference room management based on image data.

[0132] See Figure 8 , Figure 8 This is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application. Figure 8 As shown, the image recognition device 800 includes:

[0133] The first acquisition module 801 is used to acquire an image sequence of the target location, the image sequence including multiple frames of the first image;

[0134] The first extraction module 802 is used to extract features from the multi-frame first image to obtain a first feature vector, the first feature vector being used to represent motion information in the multi-frame first image;

[0135] The second extraction module 803 is used to extract features from the multi-frame first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multi-frame first image.

[0136] The recognition module 804 is used to perform object recognition on the multi-frame first image based on the first feature vector and the second feature vector.

[0137] Optionally, the image recognition device 800 may also include:

[0138] The deletion module is used to delete edge pixels of the multiple frames of the first image to generate multiple frames of the second image;

[0139] The third extraction module is used to extract features from the multi-frame second images to obtain a third feature vector, which is used to represent the motion information in the multi-frame second images.

[0140] The identification module 804 may specifically include:

[0141] A cascade unit is used to cascade the first feature vector, the second feature vector, and the third feature vector to obtain a target feature vector;

[0142] The recognition unit is used to perform object recognition on the multi-frame first image based on the target feature vector.

[0143] Optionally, the third extraction module may specifically include:

[0144] The first extraction unit is used to extract the pixel intensity of each pixel in the multiple frames of the second image in multiple preset directions using multiple preset first masks.

[0145] The first acquisition unit is used to acquire a first pixel and a second pixel in the multiple frames of the second image, wherein the first pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the second pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions.

[0146] The first construction unit is used to construct the third feature vector based on the first pixel and the second pixel.

[0147] Optionally, the first extraction module 802 may specifically include:

[0148] The second extraction unit is used to extract the pixel intensity of each pixel in the multiple frames of the first image in the multiple preset directions using a plurality of pre-set second masks;

[0149] The second acquisition unit is used to acquire the third pixel and the fourth pixel in the multi-frame first image, wherein the third pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions.

[0150] The second construction unit is used to construct the first feature vector based on the third pixel and the fourth pixel.

[0151] Optionally, the second extraction module 803 may specifically include:

[0152] The third acquisition unit is used to acquire the center pixel of the multi-frame first image and a plurality of third pixels surrounding the center pixel;

[0153] The third construction unit is used to construct the second feature vector based on the center pixel and the plurality of third pixels.

[0154] Optionally, the third building unit may specifically include:

[0155] A subunit is used to acquire multiple differences between the center pixel and the plurality of third pixels;

[0156] Construct sub-units for using a step function and the plurality of differences to construct the second feature vector.

[0157] Optionally, the image recognition device 800 may also include:

[0158] A module is created to create multiple parallel threads, which are used to acquire multiple frames of third images.

[0159] The second acquisition module is used to acquire the region of interest of the multi-frame third image, wherein the target position is located within the region of interest;

[0160] The first acquisition module 801 may specifically include:

[0161] The fourth acquisition unit is used to acquire an image sequence of the target location based on the multiple frames of the third image.

[0162] Optionally, the image recognition device 800 may also include:

[0163] The third acquisition module is used to acquire the grayscale difference between the multiple frames of the first image;

[0164] The first extraction module 802 may specifically include:

[0165] The third extraction module is used to extract features from the multiple frames of the first image when the grayscale difference is less than a preset threshold.

[0166] The image recognition device 800 can realize the embodiments of this application. Figure 1 The various processes in the method embodiments, and the ways to achieve the same beneficial effects, will not be repeated here to avoid repetition.

[0167] This application also provides a communication device. Because the principle by which the communication device solves the problem is similar to that in the embodiments of this application... Figure 1 The image recognition method shown is similar; therefore, the implementation of this communication device can be found in the implementation of the method, and repeated details will not be elaborated further. For example... Figure 9 As shown, the communication device in this embodiment includes a processor 900, configured to read a program from a memory 920 and execute the following processes:

[0168] Acquire an image sequence of the target location, the image sequence comprising multiple frames of a first image;

[0169] Feature extraction is performed on the multiple frames of the first image to obtain a first feature vector, which is used to represent the motion information in the multiple frames of the first image.

[0170] Feature extraction is performed on the multiple frames of the first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multiple frames of the first image;

[0171] Based on the first feature vector and the second feature vector, object recognition is performed on the multiple frames of the first image.

[0172] Transceiver 910 is used to receive and send data under the control of processor 900.

[0173] Among them, Figure 9 In this context, the bus architecture can include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 900) and memory (memory 920). The bus architecture can also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 910 can be multiple elements, including transmitters and transceivers, providing a unit for communicating with various other devices over a transmission medium. The processor 900 is responsible for managing the bus architecture and general processing, and the memory 920 can store data used by the processor 900 during operation.

[0174] Optionally, the processor 900 is also used to read the program from the memory 920 and perform the following steps:

[0175] Delete the edge pixels of the multi-frame first image to generate multi-frame second image;

[0176] Feature extraction is performed on the multiple frames of the second image to obtain a third feature vector, which is used to represent the motion information in the multiple frames of the second image;

[0177] The step of performing object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector includes:

[0178] The first feature vector, the second feature vector, and the third feature vector are concatenated to obtain the target feature vector;

[0179] Based on the target feature vector, object recognition is performed on the multi-frame first image.

[0180] Optionally, the step of extracting features from the multiple frames of the second image to obtain a third feature vector may include:

[0181] Using multiple pre-set first masks, the pixel intensity of each pixel in the multiple frames of the second image is extracted in multiple preset directions;

[0182] Obtain the first pixel and the second pixel in the multiple frames of the second image, wherein the first pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the second pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions.

[0183] The third feature vector is constructed based on the first pixel and the second pixel.

[0184] Optionally, the step of extracting features from the multiple frames of the first image to obtain a first feature vector may include:

[0185] Using multiple pre-set second masks, the pixel intensity of each pixel in the multiple frames of the first image is extracted in the multiple preset directions;

[0186] Obtain the third pixel and the fourth pixel in the first image of the multiple frames, wherein the third pixel is the pixel with the largest pixel intensity among the pixel intensities of the multiple preset directions, and the fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities of the multiple preset directions.

[0187] The first feature vector is constructed based on the third pixel and the fourth pixel.

[0188] Optionally, the step of extracting features from the multiple frames of the first image to obtain a second feature vector may include:

[0189] Obtain the center pixel of the first image in multiple frames, and a plurality of third pixels surrounding the center pixel;

[0190] The second feature vector is constructed based on the center pixel and the plurality of third pixels.

[0191] Optionally, constructing the second feature vector based on the center pixel and the plurality of third pixels may include:

[0192] Obtain multiple differences between the center pixel and the plurality of third pixels;

[0193] The second feature vector is constructed using a step function and the plurality of differences.

[0194] Optionally, the processor 900 is also used to read the program from the memory 920 and perform the following steps:

[0195] Multiple parallel threads are created to acquire multiple frames of third images.

[0196] Obtain the region of interest (ROI) of the multi-frame third image, wherein the target location is located within the ROI;

[0197] The acquisition of the image sequence at the target location includes:

[0198] The image sequence of the target location is obtained based on the multi-frame third image.

[0199] Optionally, the processor 900 is also used to read the program from the memory 920 and perform the following steps:

[0200] Obtain the grayscale difference between the multiple frames of the first image;

[0201] The feature extraction of the multiple frames of the first image includes:

[0202] When the grayscale difference is less than a preset threshold, feature extraction is performed on the first multi-frame image.

[0203] The communication device provided in this application embodiment can perform the above-described... Figure 1 The method embodiments shown are similar in principle and technical effect, and will not be described again here.

[0204] This application also provides a readable storage medium for storing a program, which, when executed by a processor, implements as follows: Figure 1 The various processes in the Chinese method embodiment can achieve the same technical effect, and will not be described again here to avoid repetition.

[0205] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0206] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can be physically included separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0207] The integrated units implemented as software functional units described above can be stored in a computer-readable storage medium. These software functional units, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute some steps of the transmission and reception methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0208] The above description is the preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles described in this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An image recognition method, characterized in that, include: Acquire an image sequence of the target location, the image sequence comprising multiple frames of a first image; Feature extraction is performed on the multiple frames of the first image to obtain a first feature vector, which is used to represent the motion information in the multiple frames of the first image. Feature extraction is performed on the multiple frames of the first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multiple frames of the first image; Based on the first feature vector and the second feature vector, object recognition is performed on the multiple frames of the first image; The step of extracting features from the multiple frames of the first image to obtain a first feature vector includes: Using multiple pre-set second masks, the pixel intensity of each pixel in the multiple frames of the first image is extracted in multiple preset directions; Obtain the third pixel and the fourth pixel in the first image of the multiple frames, wherein the third pixel is the pixel with the largest pixel intensity among the pixel intensities of the multiple preset directions, and the fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities of the multiple preset directions. The first feature vector is constructed based on the third pixel and the fourth pixel.

2. The method as described in claim 1, characterized in that, Before performing object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector, the method further includes: Delete the edge pixels of the multi-frame first image to generate multi-frame second image; Feature extraction is performed on the multiple frames of the second image to obtain a third feature vector, which is used to represent the motion information in the multiple frames of the second image; The step of performing object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector includes: The first feature vector, the second feature vector, and the third feature vector are concatenated to obtain the target feature vector; Based on the target feature vector, object recognition is performed on the multi-frame first image.

3. The method as described in claim 2, characterized in that, The step of extracting features from the multiple frames of the second image to obtain a third feature vector includes: Using multiple pre-set first masks, the pixel intensity of each pixel in the multiple frames of the second image is extracted in multiple preset directions; Obtain the first pixel and the second pixel in the multiple frames of the second image, wherein the first pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the second pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions. The third feature vector is constructed based on the first pixel and the second pixel.

4. The method as described in claim 1, characterized in that, The step of extracting features from the multiple frames of the first image to obtain a second feature vector includes: Obtain the center pixel of the first image in multiple frames, and a plurality of third pixels surrounding the center pixel; The second feature vector is constructed based on the center pixel and the plurality of third pixels.

5. The method as described in claim 4, characterized in that, The construction of the second feature vector based on the center pixel and the plurality of third pixels includes: Obtain multiple differences between the center pixel and the plurality of third pixels; The second feature vector is constructed using a step function and the plurality of differences.

6. The method as described in claim 1, characterized in that, Before acquiring the image sequence of the target location, the method further includes: Multiple parallel threads are created to acquire multiple frames of third images. Obtain the region of interest (ROI) of the multi-frame third image, wherein the target location is located within the ROI; The acquisition of the image sequence at the target location includes: The image sequence of the target location is obtained based on the multi-frame third image.

7. The method according to any one of claims 1 to 6, characterized in that, Before performing feature extraction on the multiple frames of the first image, the method further includes: Obtain the grayscale difference between the multiple frames of the first image; The feature extraction of the multiple frames of the first image includes: When the grayscale difference is less than a preset threshold, feature extraction is performed on the first multi-frame image.

8. An image recognition device, characterized in that, include: The first acquisition module is used to acquire an image sequence of the target location, the image sequence including multiple frames of the first image; A first extraction module is used to extract features from the multi-frame first image to obtain a first feature vector, the first feature vector being used to represent motion information in the multi-frame first image; The second extraction module is used to extract features from the multiple frames of the first image to obtain a second feature vector, which is used to represent the temporal and spatial information of the multiple frames of the first image. The recognition module is used to perform object recognition on the multiple frames of the first image based on the first feature vector and the second feature vector; The first extraction module specifically includes: The second extraction unit is used to extract the pixel intensity of each pixel in the multiple frames of the first image in multiple preset directions using multiple preset second masks; The second acquisition unit is used to acquire the third pixel and the fourth pixel in the multi-frame first image, wherein the third pixel is the pixel with the largest pixel intensity among the pixel intensities in the multiple preset directions, and the fourth pixel is the pixel with the smallest pixel intensity among the pixel intensities in the multiple preset directions. The second construction unit is used to construct the first feature vector based on the third pixel and the fourth pixel.

9. A communication device, comprising: A transceiver, a memory, a processor, and a program stored in the memory and executable on the processor; characterized in that, The processor is configured to read a program from memory to implement the steps of the method as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, A program is stored on the readable storage medium, which, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Target object determination method and device, storage medium and processor

    CN109886130A

  • Face recognition method, device and equipment and computer readable storage medium

    CN110363081A