Method, device, electronic device and computer-readable medium for generating regional information

By generating a set of feature and spatial vector information and inputting them into an embedded neural network model, the problem of excessive computer video memory and internal memory usage is solved, and the image processing efficiency and the accuracy of spatial vector information are improved.

CN114723933BActive Publication Date: 2025-09-16CHONGQING ZHONGXING MICRO ARTIFICIAL INTELLIGENCE CHIP TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011506435.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-18
Publication Date
2025-09-16
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

In the prior art, during the region information generation process, computers directly process images, resulting in excessive usage of video memory and internal memory, reducing image processing efficiency, and generating inaccurate spatial vector information.

Method used

By acquiring an image of the target object, a pre-deployed feature extraction model is used to generate a feature vector set, and a spatial vector information set is generated based on the feature vector set. Finally, it is input into a pre-trained embedded neural network model to generate the regional information of the target object.

Benefits of technology

It reduces the computer's video memory and memory usage, improves image processing efficiency, and increases the accuracy of spatial vector information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114723933B_ABST
    Figure CN114723933B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method, apparatus, electronic device, and computer-readable medium for generating region information. A specific implementation of the method includes: acquiring an image of a target object as a target image; generating a feature vector set based on the target image and a pre-deployed feature extraction model; generating a spatial vector information set based on the feature vector set; and inputting the spatial vector information set into a pre-trained embedded neural network model to obtain region information of the target object. This implementation can reduce the computer's video memory and internal memory usage, thereby reducing the computer's image processing time and, in turn, improving the computer's image processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method, device, electronic device, and computer-readable medium for generating region information. Background Art

[0002] Region information generation is a fundamental technology in the field of computer vision. Currently, methods for region information generation typically process images and extract feature information, and then process the extracted feature information to generate region information.

[0003] However, when using the above method, the following technical problems often occur:

[0004] First, because the computer processes the image directly, it occupies a large amount of computer video memory and internal memory, which increases the computer's image processing time and reduces the computer's image processing efficiency.

[0005] Second, since the influencing factors of the generation of spatial vector information are not comprehensively considered, the generated spatial vector information is not accurate enough. Summary of the Invention

[0006] The content of this disclosure is used to briefly introduce concepts that will be described in detail in the detailed description section below. The content of this disclosure is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0007] Some embodiments of the present disclosure provide a method, apparatus, electronic device, and computer-readable medium for generating region information to solve one or more of the technical problems mentioned in the above background technology section.

[0008] In a first aspect, some embodiments of the present disclosure provide a method for generating region information, the method comprising: acquiring an image of a target object as a target image; generating a feature vector set based on the target image and a pre-deployed feature extraction model; generating a spatial vector information set based on the feature vector set, wherein the spatial vector information represents the spatial structure between feature vectors in the feature vector set corresponding to the spatial vector information; and inputting the spatial vector information set into a pre-trained embedded neural network model to obtain region information of the target object.

[0009] In a second aspect, some embodiments of the present disclosure provide a region information generation device, which includes: an acquisition unit, configured to acquire an image of a target object as a target image; a first generation unit, configured to generate a feature vector set based on the above target image and a pre-deployed feature extraction model; a second generation unit, configured to generate a spatial vector information set based on the above feature vector set, wherein the above spatial vector information represents the spatial structure between feature vectors in the above feature vector set corresponding to the above spatial vector information; and an input unit, configured to input the above spatial vector information set into a pre-trained embedded neural network model to obtain region information of the target object.

[0010] In a third aspect, some embodiments of the present disclosure provide an electronic device comprising: one or more processors; a storage device on which one or more programs are stored, and when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in the first aspect.

[0011] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in the first aspect is implemented.

[0012] The aforementioned embodiments of the present disclosure have the following beneficial effects: Region information of a target object is obtained through the region information generation methods of some embodiments of the present disclosure, thereby improving the computer's image processing efficiency. Specifically, the reduced computer image processing efficiency is caused by the computer directly processing the image, which occupies a large amount of the computer's video memory and internal memory, increasing the computer's image processing time. Based on this, the region information generation methods of some embodiments of the present disclosure first acquire an image of the target object as a target image. By acquiring an image including the target object, processing of the target object is facilitated. Second, a set of feature vectors is generated based on the target image and a pre-deployed feature extraction model. This eliminates the need for subsequent image processing to process the entire image, but rather the extracted feature vectors. This reduces the space occupied by the computer's video memory and internal memory, thereby increasing the available space. Next, a set of spatial vector information is generated based on the set of feature vectors, where the spatial vector information represents the spatial structure between the individual feature vectors in the set. The obtained set of feature vectors can then be converted into spatial vector information, thereby constructing the spatial structure of the set of feature vectors. Furthermore, the spatial relationship between each feature vector can be determined. Finally, the above spatial vector information set is input into a pre-trained embedded neural network model to obtain the target object's regional information. By extracting feature vectors from the target image and obtaining the target object's regional information using the pre-trained embedded neural network model, the computer's video memory and internal memory usage are reduced, shortening the computer's image processing time. This, in turn, improves the computer's image processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that components and elements are not necessarily drawn to scale.

[0014] Figure 1 is a schematic diagram of an application scenario of the region information generation method according to some embodiments of the present disclosure;

[0015] Figure 2 is a flow chart of some embodiments of the method for generating region information according to the present disclosure;

[0016] Figure 3 is a schematic structural diagram of some embodiments of the region information generating device according to the present disclosure;

[0017] Figure 43 is a schematic structural diagram of an electronic device according to the region information generating method disclosed herein. DETAILED DESCRIPTION

[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein. On the contrary, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.

[0019] It should also be noted that, for ease of description, only the parts related to the invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present disclosure may be combined with each other.

[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0023] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0024] Figure 1 It is a schematic diagram of an application scenario of the region information generation method according to some embodiments of the present disclosure.

[0025] exist Figure 1 In an application scenario, computing device 101 may first obtain an image of a target object as target image 102. Then, computing device 101 may generate a feature vector set 104 based on target image 102 and a pre-deployed feature extraction model 103. Subsequently, based on feature vector set 104, a spatial vector information set 105 is generated. The spatial vector information represents the spatial structure between feature vectors in feature vector set 104 that correspond to the spatial vector information. Finally, spatial vector information set 105 is input into a pre-trained embedded neural network model 106 to obtain target object region information 107.

[0026] It should be noted that the computing device 101 can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed in the hardware devices listed above. It can be implemented as multiple software or software modules for providing distributed services, or as a single software or software module. No specific limitations are given here.

[0027] It should be understood that Figure 1 The number of computing devices in the embodiment is merely illustrative. Any number of computing devices may be provided according to implementation requirements.

[0028] Continue to refer Figure 2 , shows a process 200 of some embodiments of the method for generating region information according to the present disclosure. The method for generating region information includes the following steps:

[0029] Step 201: Acquire an image of a target object as a target image.

[0030] In some embodiments, the execution subject of the region information generation method (eg Figure 1 The computing device 101 shown can obtain an image of a target object in a terminal device via a wired connection or a wireless connection, and use the image as the target image. The image of the target object can include the target object and a background image. The background image can include an image of the target image that does not include the target object.

[0031] Step 202: Generate a feature vector set based on the target image and a pre-deployed feature extraction model.

[0032] In some embodiments, the execution entity may generate a feature vector set based on the target image and a pre-deployed feature extraction model. The feature extraction model may include a radar feature extraction model, a binocular camera-specific feature extraction model, and a spatial feature extraction model. The radar feature extraction model may be a model for analyzing the pixel values ​​of the target object. The binocular camera-specific feature extraction model may be a model for processing images. The spatial feature extraction model may be a model for establishing spatial vector information. The spatial vector information may characterize the spatial structure between the feature vectors corresponding to the spatial vector information in the feature vector set. The feature vector may be a feature vector extracted from the target image. The extraction may be performed by extracting feature vectors using CNN (Convolutional Neural Networks) and RNN (Recursive Neural Network).

[0033] The above-mentioned radar feature extraction model, binocular camera-specific feature extraction model and spatial feature extraction model can be trained by convolutional neural networks and deep neural networks.

[0034] In some optional implementations of some embodiments, the execution entity may generate a feature vector set by performing the following steps:

[0035] The first step is to perform grayscale processing on the target image to obtain the grayscale value of each pixel in the target image and the grayscale value of each background pixel in the background image corresponding to the target image. The grayscale processing can be performed using a component method, a maximum method, an average method, or a weighted average method to obtain the grayscale value of each pixel in the target image and the grayscale value of each background pixel in the background image corresponding to the target image.

[0036] In the second step, a difference is generated based on the grayscale value of each pixel in the target image and the grayscale value of the background pixel corresponding to the pixel in the background image corresponding to the target image, thereby obtaining a difference value set.

[0037] As an example, the grayscale value of a pixel in the target image may be 255. The grayscale value of the background pixel corresponding to the target image in the background image may be 55. The generated difference value is 200.

[0038] In the third step, a difference detection process is performed on each difference value in the difference value set to generate a target point, thereby obtaining a target point set. The difference detection process may include determining a pixel point corresponding to a difference value greater than 200 as a target point. The target point may be a pixel point in the target object.

[0039] The fourth step is to aggregate the target point set to generate a feature vector set. The aggregation process may be to classify the dispersed target point set according to the relationship between the target points and connect the target points of the same category to generate the feature vector.

[0040] Optionally, the execution entity may generate the difference by following the steps below:

[0041] In the first step, in response to the difference between the pixel point and the background pixel point satisfying a first predetermined condition, a first predetermined threshold is determined as the difference value. The first predetermined condition may be that the difference between the pixel point and the background pixel point is greater than 200. The first predetermined threshold may be 250.

[0042] In a second step, in response to the difference between the pixel point and the background pixel point satisfying a second predetermined condition, a second predetermined threshold is determined as the difference value. The second predetermined condition may be that the difference between the pixel point and the background pixel point is less than or equal to 200. The second predetermined threshold value may be 0.

[0043] Optionally, the execution entity may generate a feature vector set by performing the following steps:

[0044] The first step is to generate coordinate information of a center point of a target object's connected region based on the target point set. The center point of the target connected region may be the center point of a region including the target object. The coordinate information of the center point of the target object's connected region may be the coordinate information of the center point of a region including the target object.

[0045] The second step is to generate parameter information of the target object connected area based on the coordinate information of the target object connected area and the center point of the target object connected area. The execution entity may combine the coordinates of the center point corresponding to the coordinate information of the center point with at least one key coordinate point on the boundary of the target object connected area to generate the parameter information of the target object connected area. The boundary of the target object connected area may be generated by sequentially connecting the coordinates of the at least one key point.

[0046] As an example, the center point coordinates may be [1, 2]. The at least one key point coordinates may be (1, 1), (2, 3), (3, 3). The parameter information may be (2, 3), (1, 1), (3, 3).

[0047] The third step is to generate a set of feature vectors based on the above parameter information, wherein the feature vectors may be generated by taking a vector formed by any two coordinates included in the above parameter information as the feature vector.

[0048] As an example, the above-mentioned feature vector set may be {(1, 1)→(2, 3), (2, 3)→(3, 3), (3, 3)→(1, 1)}.

[0049] Step 203: Generate a space vector information set based on the feature vector set.

[0050] In some embodiments, the execution entity may generate a set of spatial vector information based on the feature vector set. The spatial vector information may represent the spatial structure between the feature vectors corresponding to the spatial vector information in the feature vector set. Generating the spatial vector information may involve performing a coordinate system transformation on the two coordinates included in each feature vector, converting them to a world coordinate system, generating two spatial coordinates (so that the two-dimensional coordinates become three-dimensional coordinates), and then generating the spatial vector information based on the two spatial coordinates.

[0051] As an example, the above-mentioned space vector information may be (1, 1)→(2, 3)→(3, 3)→(1, 1).

[0052] In some optional implementations of some embodiments, generating a spatial vector information set based on the feature vector set may include the following steps:

[0053] Based on the above eigenvector set, the following formula is used to generate the spatial vector information set:

[0054]

[0055] Among them, M t Represents spatial vector information. V ct represents the coordinate information of the center point of the connected area of ​​the target image. c represents the center point of the connected area of ​​the target image. t represents the frame number of the target image in the acquired image set. α represents the frame interval value. V c (t+α) represents the coordinate information of the center point of the connected area corresponding to the image of the t+αth frame. express to V ct The vector formed by to V c(t+α) The angle between the vectors formed. Indicates the The coordinate information of the center point of the connected region corresponding to the image of the frame. M represents the above-mentioned spatial vector information set. n represents the number of spatial vector information included in the above-mentioned spatial vector information set.

[0056] The above formula, as an inventive feature of an embodiment of the present disclosure, resolves the second technical problem mentioned in the background art: "The generated spatial vector information is inaccurate due to a failure to comprehensively consider the factors influencing the generation of spatial vector information." Factors that often lead to inaccurate generated spatial vector information are as follows: Failure to comprehensively consider the factors influencing the generation of spatial vector information. Resolving these factors can improve the accuracy of the generated spatial vector information. To achieve this, the above formula incorporates the connected regions of the target image to preliminarily determine the approximate region of the target object. Considering that the center point can determine the position of the target object, the present disclosure incorporates the coordinates of the center point of the connected region of the target image. This allows estimation of the approximate region of the target object in the target image. To determine how the center point of the connected region of the target image changes between different frames, the difference between the center point of the connected region of the target image and the center point of the connected region of the target image after a certain number of frames has passed is divided by the number of frames passed. This yields the change in the center point of the connected region of the image at different frame intervals. Considering the differences in the spatial positions of different eigenvectors, the vector angle is introduced. This allows precise determination of the position of each spatial vector. Since the influencing factors of the space vector information are comprehensively considered, the accuracy of the generated space vector information is improved.

[0057] Step 204: input the spatial vector information set into a pre-trained embedded neural network model to obtain the region information of the target object.

[0058] In some embodiments, the execution entity may input the spatial vector information into a pre-trained embedded neural network model to obtain the target object's region information. The pre-trained embedded neural network model may be a convolutional neural network model or a recurrent neural network model. The target object's region information may be obtained by inputting the spatial vector information set into the pre-trained embedded neural network model.

[0059] Optionally, the execution entity may send the region information of the target object to an associated device for the device to perform image processing.

[0060] As an example, the execution entity may send the region information of the target object to an image processor, so that the image processor performs image processing on the region information.

[0061] The aforementioned embodiments of the present disclosure have the following beneficial effects: Region information of a target object is obtained through the region information generation methods of some embodiments of the present disclosure, thereby improving the computer's image processing efficiency. Specifically, the reduced computer image processing efficiency is caused by the computer directly processing the image, which occupies a large amount of the computer's video memory and internal memory, increasing the computer's image processing time. Based on this, the region information generation methods of some embodiments of the present disclosure first acquire an image of the target object as a target image. By acquiring an image including the target object, processing of the target object is facilitated. Second, a set of feature vectors is generated based on the target image and a pre-deployed feature extraction model. This eliminates the need for subsequent image processing to process the entire image, but rather the extracted feature vectors. This reduces the space occupied by the computer's video memory and internal memory, thereby increasing the available space. Next, a set of spatial vector information is generated based on the set of feature vectors, where the spatial vector information represents the spatial structure between the individual feature vectors in the set. The obtained set of feature vectors can then be converted into spatial vector information, thereby constructing the spatial structure of the set of feature vectors. Furthermore, the spatial relationship between each feature vector can be determined. Finally, the above spatial vector information set is input into a pre-trained embedded neural network model to obtain the target object's regional information. By extracting feature vectors from the target image and obtaining the target object's regional information using the pre-trained embedded neural network model, the computer's video memory and internal memory usage are reduced, shortening the computer's image processing time. This, in turn, improves the computer's image processing efficiency.

[0062] Further references Figure 3 As an implementation of the above methods in the above figures, the present disclosure provides some embodiments of a region information generating device. These device embodiments are similar to Figure 2 Corresponding to the above method embodiments, the device can be specifically applied to various electronic devices.

[0063] like Figure 3As shown, the region information generating device 300 of some embodiments includes: an acquisition unit 301, a first generation unit 302, a second generation unit 303 and an input unit 304. The acquisition unit 301 is configured to acquire an image of a target object as a target image. The first generation unit 302 is configured to generate a feature vector set based on the target image and a pre-deployed feature extraction model. The second generation unit 303 is configured to generate a spatial vector information set based on the feature vector set, wherein the spatial vector information represents the spatial structure between the feature vectors in the feature vector set corresponding to the spatial vector information. The input unit 304 is configured to input the spatial vector information set into a pre-trained embedded neural network model to obtain the region information of the target object.

[0064] It is understood that the units described in the device 300 are similar to those in the reference Figure 2 Therefore, the operations, features and beneficial effects described above for the method are also applicable to the device 300 and the units included therein, and will not be repeated here.

[0065] Reference below Figure 4 , which shows an electronic device suitable for implementing some embodiments of the present disclosure (e.g., Figure 1 Schematic diagram of the structure of the computing device 101)400. Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0066] like Figure 4 As shown, the electronic device 400 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of the electronic device 400 are also stored in the RAM 403. The processing device 401, the ROM 402, and the RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 404 is also connected to the bus 404.

[0067] Typically, the following devices may be connected to the I / O interface 404: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Figure 4The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead. Figure 4 Each block shown in the figure may represent one device, or may represent multiple devices as needed.

[0068] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In some such embodiments, the computer program can be downloaded and installed from a network via the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the method of some embodiments of the present disclosure are performed.

[0069] It should be noted that the computer-readable medium described in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0070] In some embodiments, the client and server can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.

[0071] The computer-readable medium may be included in the apparatus, or may exist independently and not be incorporated into the electronic device. The computer-readable medium carries one or more programs. When executed by the electronic device, the electronic device: acquires an image of a target object as a target image; generates a feature vector set based on the target image and a pre-deployed feature extraction model; generates a spatial vector information set based on the feature vector set, wherein the spatial vector information represents the spatial structure between feature vectors in the feature vector set corresponding to the spatial vector information; and inputs the spatial vector information set into a pre-trained embedded neural network model to obtain regional information of the target object.

[0072] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0073] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0074] The units described in some embodiments of the present disclosure may be implemented in software or hardware. The units described may also be provided in a processor. For example, they may be described as follows: a processor includes an acquisition unit, a first generation unit, a second generation unit, and an input unit. The names of these units do not, in some cases, constitute limitations on the units themselves. For example, the acquisition unit may also be described as a "unit for acquiring an image of a target object as a target image."

[0075] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0076] The above description is only an illustration of some preferred embodiments of the present disclosure and the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but should also cover other technical solutions formed by any combination of the above-mentioned technical features or their equivalent features without departing from the above-mentioned inventive concept. For example, the above-mentioned features are replaced with (but not limited to) technical features with similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A method for generating region information, comprising: acquiring an image of a target object as a target image; generating a set of feature vectors based on the target image and a pre-deployed feature extraction model; generating a spatial vector information set according to the feature vector set, wherein the spatial vector information represents a spatial structure between feature vectors in the feature vector set corresponding to the spatial vector information; Inputting the spatial vector information set into a pre-trained embedded neural network model to obtain region information of the target object; The step of generating a spatial vector information set based on the feature vector set includes: Based on the feature vector set, a spatial vector information set is generated using the following formula: Among them, M t Represents space vector information, V ct represents the coordinate information of the center point of the target image connected area, the target image connected area is the approximate area of ​​the target object, c represents the center point of the target image connected area, t represents the frame number of the target image in the acquired image set, α represents the frame interval value, V c(t+α) Represents the coordinate information of the center point of the connected area corresponding to the image of the t+αth frame, express to V ct The vector formed by to V c(t+α) The angle between the vectors formed, Indicates the The coordinate information of the center point of the connected area corresponding to the image of the frame, M represents the spatial vector information set, and n represents the number of spatial vector information included in the spatial vector information set.

2. The method according to claim 1, wherein The method further comprises: The area information of the target object is sent to an associated device for the device to perform image processing.

3. The method according to claim 2, wherein: The generating a feature vector set based on the target image and a pre-deployed feature extraction model includes: Performing image grayscale processing on the target image to obtain the grayscale value of each pixel in the target image and the grayscale value of each background pixel in the background image corresponding to the target image; generating a difference based on the grayscale value of each pixel in the target image and the grayscale value of a background pixel corresponding to the pixel in the background image corresponding to the target image, to obtain a difference value set; performing a difference detection process on each difference in the difference set to generate a target point, thereby obtaining a target point set; Aggregation processing is performed on the target point set to generate a feature vector set.

4. The method according to claim 3, wherein: Generating a difference value based on the grayscale value of each pixel in the target image and the grayscale value of a background pixel corresponding to the pixel in a background image corresponding to the target image includes: In response to a difference between the pixel point and the background pixel point satisfying a first predetermined condition, determining a first predetermined threshold as the difference value; In response to the difference between the pixel point and the background pixel point satisfying a second predetermined condition, a second predetermined threshold is determined as the difference value.

5. The method according to claim 4, wherein The feature extraction model includes a radar feature extraction model, a binocular camera-specific feature extraction model and a spatial feature extraction model. The radar feature extraction model is a model used to analyze motion speed and direction, the binocular camera-specific feature extraction model is a model used to process video streams, and the spatial feature extraction model is a model used to establish a spatial vector model.

6. The method according to claim 5, wherein: The aggregating the target point set to generate a feature vector set includes: Based on the target point set, generating coordinate information of a center point of a connected area of ​​a target object; generating parameter information of the target object connected area based on the coordinate information of the target object connected area and the center point of the target object connected area; Based on the parameter information, a feature vector set is generated.

7. A region information generating device, comprising: an acquisition unit configured to acquire an image of a target object as a target image; A first generating unit is configured to generate a feature vector set based on the target image and a pre-deployed feature extraction model; a second generating unit configured to generate a spatial vector information set based on the feature vector set, wherein the spatial vector information represents a spatial structure between feature vectors in the feature vector set corresponding to the spatial vector information; an input unit configured to input the spatial vector information set into a pre-trained embedded neural network model to obtain region information of the target object; The step of generating a spatial vector information set based on the feature vector set includes: Based on the feature vector set, a spatial vector information set is generated using the following formula: Among them, M t Represents space vector information, V ct represents the coordinate information of the center point of the target image connected area, the target image connected area is the approximate area of ​​the target object, c represents the center point of the target image connected area, t represents the frame number of the target image in the acquired image set, α represents the frame interval value, V c(t+α) Represents the coordinate information of the center point of the connected area corresponding to the image of the t+αth frame, express to V ct The vector formed by to V c(t+α) The angle between the vectors formed, Indicates the The coordinate information of the center point of the connected area corresponding to the image of the frame, M represents the spatial vector information set, and n represents the number of spatial vector information included in the spatial vector information set.

8. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Re-identification method and device for pedestrians in video image, storage medium and processor

    CN107844753A

  • Video retrieval device, video retrieval method used therefor, and program therefor

    JP2003345830A