Training method and device of object recognition network, storage medium and electronic equipment
By introducing the distribution information of a large model into the object recognition network as supervision, a small model is trained and made to converge in different hardware environments. This solves the problem of the object recognition network's inability to transfer, and improves the model's running efficiency and transferability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-26
- Publication Date
- 2026-03-17
AI Technical Summary
Existing object recognition networks cannot be migrated across different hardware environments, resulting in high hardware compatibility requirements.
By performing object recognition in the first recognition network to obtain the first recognition result, and then performing object recognition in the second recognition network, the target recognition network is trained using the positional distribution loss of the first and second recognition results, so that it reaches the convergence condition in the hardware environment. The distribution information of the large model is introduced as supervision to encourage the small model to learn from the large model.
It improves the transferability of object recognition networks across different hardware environments, reduces the difficulty of learning the ground truth of labels, and improves the running efficiency of the model.
Smart Images

Figure CN117011663B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computers, and more specifically, to a training method and apparatus for an object recognition network, a storage medium, and an electronic device. Background Technology
[0002] Currently, there are various object recognition networks that can run on computers. However, because different object recognition networks have different model sizes, they usually have certain hardware requirements for the hardware environment in which they run. For example, the computing power of in-vehicle infotainment systems is relatively limited, so they cannot run computationally intensive object recognition networks. At the same time, while smaller models can achieve some recognition functions, they cannot achieve the high-precision output of larger models.
[0003] In other words, existing object recognition models typically have high hardware compatibility requirements for the operating environment, and there are technical problems that prevent them from being migrated and applied in different hardware environments.
[0004] There is currently no effective solution to the above problems. Summary of the Invention
[0005] This invention provides a training method and apparatus, storage medium and electronic device for an object recognition network, to at least solve the technical problem that existing object recognition networks cannot be migrated in different hardware environments.
[0006] According to one aspect of the present invention, a training method for an object recognition network is provided, comprising: acquiring a set of sample images and a set of object location labels corresponding to the set of sample images, wherein the sample images in the set of sample images contain a target object to be recognized, and the object location labels in the set of object location labels are used to indicate the location labels corresponding to the positions of the target object in the sample images; performing object recognition on the sample images in a first recognition network to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the position distribution of a first candidate location set obtained after recognizing the target object; and performing object recognition on the sample images in a second recognition network to obtain... The second recognition result corresponding to the target object is used to indicate the position distribution of the second candidate position set obtained after recognizing the target object; the network structure size of the first recognition network is smaller than the network structure size of the second recognition network; a first position distribution loss between the first recognition result and the object position label, and a second position distribution loss between the first recognition result and the second recognition result are obtained; based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached the convergence condition after training the first recognition network is obtained, wherein the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold.
[0007] According to another aspect of the present invention, a training apparatus for an object recognition network is also provided, comprising: a first acquisition unit, configured to acquire a set of sample images and a set of object location labels corresponding to the set of sample images, wherein the sample images in the set of sample images contain a target object to be identified, and the object location labels in the set of object location labels are used to indicate the location labels corresponding to the positions of the target object in the sample images; a first recognition unit, configured to perform object recognition on the sample images in a first recognition network to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the positional distribution of a first candidate location set obtained after recognizing the target object; and a second recognition unit, configured to perform object recognition on the sample images in a second recognition network. The system is configured to: identify a second identification result corresponding to the target object, wherein the second identification result indicates the positional distribution of the second candidate position set obtained after identifying the target object; the network structure size of the first identification network is smaller than the network structure size of the second identification network; a second acquisition unit is configured to acquire a first positional distribution loss between the first identification result and the object position label, and a second positional distribution loss between the first identification result and the second identification result; and a training unit is configured to obtain a target identification network that has reached convergence after training the first identification network based on the first positional distribution loss and the second positional distribution loss, wherein the similarity between the identification result output by the target identification network and the identification result output by the second identification network is less than a threshold.
[0008] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the above-described object recognition network training method at runtime.
[0009] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the training method for the object recognition network described above.
[0010] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the above-described object recognition network training method through the computer program.
[0011] In this embodiment of the invention, a set of sample images and a set of object location labels corresponding to the sample image set are obtained; object recognition is performed on the sample images in a first recognition network to obtain a first recognition result; object recognition is performed on the sample images in a second recognition network to obtain a second recognition result; the network structure size of the first recognition network is smaller than that of the second recognition network; a first position distribution loss between the first recognition result and the object location labels, and a second position distribution loss between the first recognition result and the second recognition result are obtained; based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached the convergence condition after training the first recognition network is obtained. The method described in this application introduces the distribution information of a large model as supervision, enabling the small model to learn from the large model and the feature representation to move closer to the large model, reducing the difficulty of learning the ground truth labels, improving the model's operating efficiency, and thus solving the technical problem that existing object recognition networks cannot be transferred to different hardware environments. Attached Figure Description
[0012] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0013] Figure 1 This is a schematic diagram of the hardware environment for an optional object recognition network training method according to an embodiment of the present invention;
[0014] Figure 2 This is a flowchart of an optional object recognition network training method according to an embodiment of the present invention;
[0015] Figure 3 This is a schematic diagram of an optional object recognition network training method according to an embodiment of the present invention;
[0016] Figure 4 This is a schematic diagram of another optional object recognition network training method according to an embodiment of the present invention;
[0017] Figure 5 This is a schematic diagram of the structure of a training device for an optional object recognition network according to an embodiment of the present invention;
[0018] Figure 6 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0021] The following explains the terminology used in this application:
[0022] Deep learning is a branch of machine learning that is based on neural network architecture and learns representations of data. It is divided into unsupervised, semi-supervised, and fully supervised learning and has been widely used in fields such as computer vision, speech recognition, and natural language processing.
[0023] Image elements: Useful physical point information in map data images, such as traffic restriction signs, speed limit signs, traffic lights, etc.
[0024] Large models: generally refers to a class of models that require more computation and have more parameters than ordinary models.
[0025] Knowledge distillation: Further optimize model performance through knowledge transfer.
[0026] Multi-task model: includes a series of deep learning task models such as object detection and image classification.
[0027] According to one aspect of the present invention, a method for training an object recognition network is provided. As an optional implementation, the above-described object recognition network training method can be applied, but is not limited to, to applications such as... Figure 1 The training system shown is for an object recognition network consisting of server 102 and terminal device 104. Figure 1As shown, server 102 is connected to terminal device 104 via network 110. This network may include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks (LANs), metropolitan area networks (MANs), and wide area networks (WANs). The wireless network includes Bluetooth, Wi-Fi, and other networks enabling wireless communication. The terminal device may include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, PDAs, MIDs (Mobile Internet Devices), PADs, desktop computers, smart TVs, in-vehicle devices, etc. The terminal device may have a client installed, such as an object recognition client or a game client. The terminal device also includes a display, a processor, and a memory. The display can show the program interface of the object recognition client or game client, as well as the road images to be uploaded to the server. The processor can preprocess the road image files before transmission, for example, by compressing the acquired image files. The memory stores the image files to be uploaded. It is understood that after acquiring the road image to be uploaded in the aforementioned terminal device 104, the terminal device 104 can send the road image to the server 102 via the network 110. Upon receiving the road image, the server 102 performs image annotation based on the road image uploaded by the terminal device 104 and uses it for training object recognition images. The terminal device 104 can receive the trained object recognition network returned by the server 102 via the network 110. The server 102 can be a single server, a server cluster consisting of multiple servers, or a cloud server. The aforementioned server includes a database and a processing engine. The database may include labeled samples for network training; the processing engine is used to execute the aforementioned network training process.
[0028] According to one aspect of the present invention, the training system for the object recognition network described above may further perform the following steps: Server 102 performs steps S102 to S110 to obtain a set of sample images and a set of object location labels corresponding to the set of sample images, wherein the sample images in the set of sample images contain a target object to be recognized, and the object location labels in the set of object location labels are used to indicate the location label corresponding to the position of the target object in the sample image; object recognition is performed on the sample images in a first recognition network to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the location distribution of the first candidate location set obtained after recognizing the target object; object recognition is performed on the sample images in a second recognition network to obtain a second recognition result corresponding to the target object, wherein the second recognition... The result is used to indicate the positional distribution of the second candidate position set obtained after the target object is identified; the network structure size of the first identification network is smaller than that of the second identification network; the first positional distribution loss between the first identification result and the object position label, and the second positional distribution loss between the first identification result and the second identification result are obtained; based on the first positional distribution loss and the second positional distribution loss, a target identification network that has reached the convergence condition after training the first identification network is obtained, wherein the similarity between the identification result output by the target identification network and the identification result output by the second identification network is less than a threshold; then, the server 102 executes step S112, and sends the target identification network to the terminal device 104 through the network 110; finally, the terminal device 104 executes step S114, and uses the target identification network to perform object identification.
[0029] In this embodiment of the invention, a set of sample images and a set of object location labels corresponding to the sample image set are obtained; object recognition is performed on the sample images in a first recognition network to obtain a first recognition result; object recognition is performed on the sample images in a second recognition network to obtain a second recognition result; the network structure size of the first recognition network is smaller than that of the second recognition network; a first position distribution loss between the first recognition result and the object location labels, and a second position distribution loss between the first recognition result and the second recognition result are obtained; based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached the convergence condition after training the first recognition network is obtained. The method described in this application introduces the distribution information of a large model as supervision, enabling the small model to learn from the large model and the feature representation to move closer to the large model, reducing the difficulty of learning the ground truth labels, improving the model's operating efficiency, and thus solving the technical problem that existing object recognition networks cannot be transferred to different hardware environments.
[0030] The above is merely an example, and no limitation is made in this embodiment.
[0031] As an optional implementation method, such as Figure 2 As shown, the training method for the above object recognition network includes the following steps:
[0032] S202, Obtain the sample image set and the object location label set corresponding to the sample image set;
[0033] Among them, the sample images in the sample image set contain the target objects to be identified, and the object location labels in the object location label set are used to indicate the location labels corresponding to the positions of the target objects in the sample images;
[0034] It is understood that the aforementioned sample image set may include multiple images, specifically road images acquired through vehicle-mounted imaging equipment. Specifically, it may include... Figure 3 As shown, Figure 3 A specific sample image is shown, in which multiple objects are labeled with lines and boxes. For example, the first object 301, labeled with a box, can be a road arrow object; the second object 302, labeled with a box, is another road arrow object; and the third object 303, labeled with lines, is a road edge object. It is understood that the sample image set includes multiple objects such as... Figure 3 The sample images shown each contain multiple identified and labeled objects.
[0035] Understandably, the specific locations of different objects in an image can be marked using annotation lines and boxes. Taking the first object 301 as an example, its location in the image can be indicated by the coordinates of its upper left and lower right corners in its corresponding annotation box. For example, a box with coordinates (30pt, 50pt) and (100pt, 5pt) as its diagonal points could indicate the location of the first object 301 in the image. Taking the third object 303 as an example, its location in the image can be indicated by the coordinates of the start and end points of its corresponding annotation line. For example, the location of an annotation line starting at coordinates (230pt, 60pt) and ending at coordinates (280pt, 2pt) could indicate the location of the third object 303 in the image.
[0036] S204, In the first recognition network, object recognition is performed on the sample image to obtain a first recognition result corresponding to the target object;
[0037] The first identification result is used to indicate the location distribution of the first candidate location set obtained after identifying the target object;
[0038] It should be noted that the aforementioned first identification network is an identification network that can operate in an in-vehicle system. Through this first identification network, [the system can identify...] Figure 3The image shown is used for object recognition, and a first recognition result corresponding to the identified target object can be obtained. Assume the above first recognition network... Figure 3 When performing object recognition, if the identified target object includes the first object 301, the position of the first object 301 in the image can be output first. Assuming that the first recognition network can output a box with the coordinates (25pt, 40pt) and the coordinates (110pt, 10pt) as the diagonal to indicate the position of the first object 301 in the image.
[0039] Furthermore, the first recognition network described above can further output the corresponding position distribution based on the aforementioned coordinate points, serving as the first recognition result. This position distribution can be a distribution for each recognition point, and can be a mathematical distribution, such as a normal distribution. Optionally, the position distribution can also be multiple discrete results obtained during the model's training process. For example, it can be composed of multiple candidate position points and their corresponding probability values. For instance, for coordinate points (25pt, 40pt), the output coordinate point distribution could be (10pt, 40pt, 10%), (25pt, 35pt, 20%), (25pt, 40pt, 50%), (35pt, 50pt, 20%); for coordinate points (110pt, 10pt), the output coordinate point distribution could be (100pt, 10pt, 10%), (110pt, 15pt, 20%), (110pt, 10pt, 50%), (125pt, 5pt, 20%).
[0040] The above method for obtaining the location distribution of the first candidate location set is only an example and does not limit the actual method of obtaining it.
[0041] S206, In the second recognition network, object recognition is performed on the sample image to obtain a second recognition result corresponding to the target object;
[0042] The second recognition result is used to indicate the positional distribution of the second candidate position set obtained after recognizing the target object; the network structure size of the first recognition network is smaller than the network structure size of the second recognition network.
[0043] It should be noted that the second recognition network described above can identify the target object, output coordinate information indicating the target object's location, and further generate the second recognition result based on the coordinate information. This second recognition network can run on a computer terminal, but its model size and computing power requirements are higher than those of the first recognition network, therefore it cannot be directly run on hardware devices with weak computing power.
[0044] S208, obtain the first position distribution loss between the first recognition result and the object position label, and the second position distribution loss between the first recognition result and the second recognition result;
[0045] S210, based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached convergence after training the first recognition network is obtained. Specifically, the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold.
[0046] The above embodiments of this application involve obtaining a set of sample images and a set of object location labels corresponding to the sample image set; performing object recognition on the sample images in a first recognition network to obtain a first recognition result; performing object recognition on the sample images in a second recognition network to obtain a second recognition result; the network structure size of the first recognition network is smaller than that of the second recognition network; obtaining a first position distribution loss between the first recognition result and the object location labels, and a second position distribution loss between the first recognition result and the second recognition result; and obtaining a target recognition network that has reached convergence after training the first recognition network based on the first position distribution loss and the second position distribution loss. The method of this application introduces the distribution information of a large model as supervision, enabling the small model to learn from the large model and the feature representation to move closer to the large model, reducing the difficulty of learning the ground truth labels, improving the model's operating efficiency, and thus solving the technical problem that existing object recognition networks cannot be transferred to different hardware environments.
[0047] As an optional implementation, the above-mentioned first location distribution loss between the first recognition result and the object location label, and the second location distribution loss between the first recognition result and the second recognition result include:
[0048] S1, obtain a first numerical distribution sequence corresponding to the first recognition result and a second numerical distribution sequence corresponding to the second recognition result;
[0049] S2, retrieve the sequence of object position coordinate values included in the object position label;
[0050] S3, determine the first position distribution loss using the first numerical distribution sequence and the object position coordinate value sequence;
[0051] S4. Use the first numerical distribution sequence and the second numerical distribution sequence to determine the second location distribution loss.
[0052] The following provides a further description of an optional method for extracting the aforementioned numerical distribution sequence. In the above embodiments of this application, it is assumed that the first identification result is the first identification network pair. Figure 3The identification results of the first object 301 include the first coordinate point (25pt, 40pt) and the second coordinate point (110pt, 10pt), whose corresponding position distributions are (10pt, 40pt, 10%), (25pt, 35pt, 20%), (25pt, 40pt, 50%), (35pt, 50pt, 20%), and (100pt, 10pt, 10%), (110pt, 15pt, 20%), (110pt, 10pt, 50%), (125pt, 50pt, 10pt). The coordinate values of each of the above-mentioned locations (10pt, 40pt, 25pt, 35pt, 25pt, 40pt, 35pt, 50pt, 35pt, 50pt, 100pt, 10pt, 110pt, 15pt, 110pt, 10pt, 125pt, 5pt) are then extracted and combined to obtain the first numerical distribution sequence (10pt, 40pt, 25pt, 35pt, 25pt, 40pt, 35pt, 50pt, 35pt, 50pt, 100pt, 10pt, 110pt, 15pt, 110pt, 10pt, 125pt, 5pt). Similarly, given the two output coordinates of the second recognition network, the corresponding second numerical distribution sequence can be obtained based on the output coordinates. It should be noted that the above methods for obtaining the first and second numerical distribution sequences are merely examples and do not limit the actual process of obtaining the numerical distributions.
[0053] Through the above-described embodiments of this application, a first numerical distribution sequence corresponding to the first recognition result and a second numerical distribution sequence corresponding to the second recognition result are obtained; an object position coordinate value sequence included in the object position label is obtained; a first position distribution loss is determined using the first numerical distribution sequence and the object position coordinate value sequence; and a second position distribution loss is determined using the first numerical distribution sequence and the second numerical distribution sequence. Thus, the loss is calculated based on the numerical distribution sequence output by the first network and the numerical distribution sequence output by the second network, thereby enabling the first network to learn from the second network to a higher degree and improving the efficiency of model transfer.
[0054] As an optional implementation, the above-described method of determining the second positional distribution loss using the first and second numerical distribution sequences includes: traversing the first and second numerical distribution sequences and performing the following steps until the traversal is complete:
[0055] S1, obtain the i-th first distribution value of the first numerical distribution sequence and the i-th second distribution value of the second numerical distribution sequence, where i is a positive integer and the i-th first distribution value corresponds to the i-th second distribution value;
[0056] S2, input the i-th second distribution value and the temperature coefficient into the normalized exponential function to obtain the i-th reference second distribution value, where the temperature coefficient is used to indicate the compactness of the simulation based on the second value distribution sequence;
[0057] S4, obtain the logarithm of the ratio of the i-th first distribution value to the i-th reference second distribution value as the weight coefficient of the i-th value;
[0058] S4, obtain the ratio of the i-th numerical weight coefficient to the i-th first distribution value as the i-th loss value;
[0059] S5, after the traversal of the first and second numerical distribution sequences is completed, obtain the sum of N loss values as the second positional distribution loss, where N is the number of loss values obtained during the traversal.
[0060] In this embodiment, the method for obtaining the numerical distribution sequence based on the position coordinates can be to concatenate the distribution sequences corresponding to each coordinate value to obtain the aforementioned numerical distribution sequence. For example, the first recognition network has recognized the target object and obtained four coordinate points (x1, y1), (x2, y2), (x3, y3), and (x4, y4). It can be understood that the four coordinate points respectively indicate the coordinates of the four vertices of the annotation box that marks the target object. Further, the distribution of the above four coordinate points can be obtained by first obtaining the distribution of x1 (a1, a2, a3, a4), then obtaining the distribution of y1 (b1, b2, b3, b4), further obtaining the distribution of x2 (a5, a6, a7, a8)... and so on, and then concatenating the distribution of each coordinate value to obtain a distribution sequence composed of 32 values. Correspondingly, when obtaining the four coordinate points obtained by the second recognition network from recognizing the target object, the distribution sequence corresponding to the four points can also be obtained.
[0061] Given the first and second numerical distribution sequences described above, the second positional distribution loss is calculated using the following formula:
[0062]
[0063] Where F_D_S indicates the i-th first distribution value, F_D_T is the i-th second distribution value, and τ is the temperature coefficient, representing the degree of compactness of the distribution simulation during training, which is generally taken as 10.
[0064] Through the above-described embodiments of this application, the i-th first distribution value of the first numerical distribution sequence and the i-th second distribution value of the second numerical distribution sequence are obtained, where i is a positive integer, and the i-th first distribution value corresponds to the i-th second distribution value; the i-th second distribution value and the temperature coefficient are input into a normalized exponential function to obtain the i-th reference second distribution value, where the temperature coefficient is used to indicate the compactness of the simulation based on the second numerical distribution sequence; the logarithm of the ratio of the i-th first distribution value to the i-th reference second distribution value is obtained as the i-th numerical weight coefficient; the i-th numerical weight coefficient is obtained... The ratio of the number to the i-th value of the first distribution is used as the i-th loss value. After traversing the first and second numerical distribution sequences, the sum of N loss values is obtained as the second positional distribution loss, where N is the number of loss values obtained during the traversal. Thus, the second positional distribution loss is determined by using the first and second numerical distribution sequences, thereby optimizing the similarity of the output probability distribution. This transforms the optimization objective into simply moving closer to the larger model in terms of probability distribution. Different loss functions are designed for different tasks, which reduces the learning difficulty to some extent and makes the optimization objective more scientific.
[0065] As an optional implementation, the above-described method of determining the first position distribution loss using the first numerical distribution sequence and the object position coordinate value sequence includes:
[0066] S1, obtain the target regression loss function and the target classification loss function;
[0067] S2, calculate the regression loss value of the first numerical distribution sequence and the object location coordinate value sequence using the target regression loss function, and calculate the first numerical distribution sequence and the object location coordinate value sequence using the classification loss function;
[0068] S3, obtain the weighted sum of the regression loss value and the classification loss value as the first position distribution loss.
[0069] Specifically, the calculation method used in this embodiment is as follows:
[0070] L1 = L reg +L cls
[0071] Among them, L reg As a regression loss function, L clsThis is a classification loss function. The regression loss function and classification loss function mentioned above can be selected according to actual needs. For example, the regression loss function can be one of the following: mean squared error, squared loss, mean absolute loss, Huber loss (smoothed mean absolute error), Log-Cosh loss, quantile loss, etc.; the classification loss function can be one of the following: 0-1 loss, log function loss, cross-entropy loss, etc.
[0072] Furthermore, the aforementioned L reg It is the regression loss value L obtained by using the first numerical distribution sequence and the sequence of coordinate values of the object's location. cls The loss can be calculated by comparing the sequence of object position coordinates output by the first recognition network with the sequence of object position coordinates in the annotation results.
[0073] Through the above-described embodiments of this application, a target regression loss function and a target classification loss function are obtained; the regression loss value of the first numerical distribution sequence and the target location coordinate value sequence is calculated using the target regression loss function, and the first numerical distribution sequence and the target location coordinate value sequence are calculated using the classification loss function; the weighted sum of the regression loss value and the classification loss value is obtained as the first location distribution loss, and the first recognition network is trained using the accurate annotation results, thereby improving the recognition accuracy of the first recognition network.
[0074] As an optional implementation, the above-described object recognition of the sample image in the first recognition network to obtain a first recognition result corresponding to the target object and object recognition of the sample image in the second recognition network to obtain a second recognition result corresponding to the target object include:
[0075] The first recognition sub-network is used to perform object recognition on the sample image to obtain a first object position coordinate value sequence; the first object position coordinate value sequence is input into the mapping sub-network to obtain a first numerical distribution sequence;
[0076] The sample image is used to perform object recognition using the second recognition sub-network to obtain a second object position coordinate value sequence; the second object position coordinate value sequence is input into the mapping sub-network to obtain a second numerical distribution sequence, wherein the network parameter scale of the first recognition sub-network is smaller than that of the second recognition sub-network.
[0077] The following combination Figure 4 The above method will be described in detail. In this embodiment, the first recognition subnetwork and the second recognition subnetwork can each perform the two steps of image input and image perception, respectively.
[0078] The image input step may include: acquiring a visible image input stream of the vehicle during its driving process through a vehicle acquisition device, such as a dashcam or a pre-installed / aftermarket camera; using the video stream processing component in the SDK to sample and segment the video stream into individual video frames according to the actual frame rate; and then using a compression algorithm to process them into image format as input for the model image.
[0079] Next, the image perception step is performed. The input image is resized to a suitable size, passed through a deep convolutional backbone network, and the basic features required for the task are extracted and sent to the downstream detection and recognition model for further perception. The image perception model consists of a backbone network and multi-task sub-networks. Different multi-task perception networks share information from the shallow backbone network. It is understandable that the backbone and multi-task sub-network structures of the first and second recognition sub-networks can be different, and the backbone and multi-task sub-network structures of the second recognition sub-network are more complex than those of the first recognition sub-network, with a larger scale of intermediate parameters. Therefore, the aforementioned second recognition network is used as the teacher model, and the initial first recognition network is used as the student model.
[0080] Due to hardware limitations, the computational cost and number of parameters in the student model are restricted. Furthermore, the implementation method described in this application uses a larger model with greater parameters and computational cost as the teacher model. The optimization goal is to enable the student model to learn from the teacher model, thereby enhancing the model's robustness.
[0081] Through the above-described embodiments of this application, the output distributions of the teacher model and the student model are obtained by using a first recognition sub-network to perform object recognition on the sample image to obtain a first object position coordinate value sequence; inputting the first object position coordinate value sequence into a mapping sub-network to obtain a first numerical distribution sequence; using a second recognition sub-network to perform object recognition on the sample image to obtain a second object position coordinate value sequence; and inputting the second object position coordinate value sequence into a mapping sub-network to obtain a second numerical distribution sequence. The loss between the output distributions of the teacher model and the student model is then used to train the student model, allowing the student model to learn from the teacher model and enhancing the robustness of the model.
[0082] As an optional implementation, the target recognition network that achieves convergence after training the first recognition network, based on the first position distribution loss and the second position distribution loss, includes:
[0083] S1, obtain the current joint coefficient, where the current joint coefficient is the weight coefficient determined based on the number of training iterations performed so far, and the current joint coefficient is less than 1;
[0084] S2, determine the current reference joint coefficient based on the difference between the value 1 and the current joint coefficient;
[0085] S3, the sum of the product of the current reference joint coefficient and the first position distribution loss and the product of the current joint coefficient and the second position distribution loss is taken as the joint loss value;
[0086] S4. If the joint loss value is less than or equal to the target loss threshold, the current first recognition network is used as the target recognition network that has reached the convergence condition; if the joint loss value is greater than the target loss threshold, the current first recognition network continues to train.
[0087] As an optional implementation, obtaining the current joint coefficients includes:
[0088] S1, if the number of training iterations is less than or equal to N1, the current joint coefficient is determined as the first joint coefficient, where the first joint coefficient is a constant greater than or equal to 0, and N1 is a positive integer;
[0089] S2, if the number of training iterations is greater than N1 and less than or equal to N2, the current joint coefficient is determined as the second joint coefficient, where the second joint coefficient increases with the number of training iterations, the second joint coefficient is greater than the first joint coefficient, and N2 is a positive integer greater than N1;
[0090] S3, if the number of training iterations is greater than N2 and less than or equal to N3, the current joint coefficient is determined as the third joint coefficient, where the third joint coefficient is a constant greater than the second joint coefficient, and N3 is a positive integer greater than N2.
[0091] It is understood that, in this embodiment, after obtaining the first location distribution loss and the second location distribution loss, the joint loss value can be further obtained using the following formula:
[0092] L all = (1-αβ)L1+α*L2
[0093] Here, α is the optimization coefficient, which is a variable. As training progresses, changes in α have a certain impact on the final optimization result. The final change process of α is as follows:
[0094]
[0095] Where e represents the epoch value of the current iteration of the model.
[0096] Through the above-described embodiments of this application, different joint loss coefficients are selected according to the number of training iterations during the training process, resulting in different training tendencies at different training stages. Specifically, in the early stages of training, the output of the first recognition network can be made closer to the true value, and in the middle and later stages of training, the output of the first recognition network can be made closer to the output of the second recognition network, thereby improving the accuracy of the small model without increasing the amount of additional parameters or computation.
[0097] The following combination Figure 4 A complete process is provided to illustrate the above method.
[0098] S1, Image Input;
[0099] Specifically, the above steps are executed in the student model and the teacher model respectively. The visible image input stream of the vehicle during the current driving process is obtained through the vehicle acquisition device, the dashcam or the vehicle's built-in pre-installed / after-installed camera. The video stream is sampled and divided into video frames according to the actual frame rate used by the video stream processing component in the SDK. Then, the frames are processed into image format by the compression algorithm and used as the input of the model image.
[0100] S2, Image Perception;
[0101] The input image is resized to a suitable size, then passed through a deep convolutional backbone network to extract the basic features required for the task, and finally fed into the downstream detection and recognition model for further perception. The image perception model consists of a backbone network and multi-task sub-networks. Different multi-task perception networks share information from the shallow backbone network.
[0102] S3, Model Training;
[0103] During model training, the current model to be optimized is typically used as the student model. Due to hardware limitations, the computational cost and number of parameters in the student model are restricted. Simultaneously, this method uses a larger model with more parameters and higher computational cost as the teacher model. The goal of optimization is to enable the student model to learn from the teacher model, thereby enhancing the model's robustness.
[0104] Taking the detection task as an example, in the anchor-free method of FCOS, for each of the four vertices of a detection box, there is a distribution for each vertex to describe the specific position of the box in the image. F_D_S represents the distribution information of the student network, and F_D_T represents the position distribution information of the teacher network.
[0105] Given the location distribution information described above, the second distribution loss is calculated using the following formula:
[0106]
[0107] Where F_D_S indicates the i-th first distribution value, F_D_T is the i-th second distribution value, and τ is the temperature coefficient, representing the degree of compactness of the distribution simulation during training, which is generally taken as 10.
[0108] In the L1 branch, the original detection training loss is retained, primarily consisting of regression and classification loss functions. The specific calculation method used is as follows:
[0109] L1 = L reg +L cls
[0110] Among them, L reg As a regression loss function, L cls This is a classification loss function. The regression loss function and classification loss function mentioned above can be selected according to actual needs. For example, the regression loss function can be one of the following: mean squared error, squared loss, mean absolute loss, Huber loss (smoothed mean absolute error), Log-Cosh loss, quantile loss, etc.; the classification loss function can be one of the following: 0-1 loss, log function loss, cross-entropy loss, etc.
[0111] After obtaining the first position distribution loss and the second position distribution loss, the joint loss value can be further obtained using the following formula:
[0112] L all = (1-α)L1+α*L2
[0113] Here, α is the optimization coefficient, which is a variable. As training progresses, changes in α have a certain impact on the final optimization result. The final change process of α is as follows:
[0114]
[0115] Where e represents the epoch value of the current iteration of the model.
[0116] The above-described embodiments of this application involve: acquiring a set of sample images and a set of object location labels corresponding to the sample image set; performing object recognition on the sample images in a first recognition network to obtain a first recognition result; performing object recognition on the sample images in a second recognition network to obtain a second recognition result; the network structure size of the first recognition network is smaller than that of the second recognition network; acquiring a first positional distribution loss between the first recognition result and the object location labels, and a second positional distribution loss between the first recognition result and the second recognition result; and obtaining a target recognition network that has reached convergence after training the first recognition network based on the first positional distribution loss and the second positional distribution loss. The method of this application introduces the distribution information of a large model as supervision, enabling the small model to learn from the large model and the feature representation to move closer to the large model, reducing the difficulty of learning the ground truth labels, improving the model's operating efficiency, and thus solving the technical problem that existing object recognition networks cannot be transferred to different hardware environments.
[0117] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0118] According to another aspect of the present invention, a training apparatus for an object recognition network for implementing the above-described object recognition network training method is also provided. For example... Figure 5 As shown, the device includes:
[0119] The first acquisition unit 502 is used to acquire a sample image set and an object location label set corresponding to the sample image set. The sample images in the sample image set contain the target object to be identified, and the object location labels in the object location label set are used to indicate the location label corresponding to the position of the target object in the sample image.
[0120] The first recognition unit 504 is used to perform object recognition on the sample image in the first recognition network to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the position distribution of the first candidate position set obtained after recognizing the target object;
[0121] The second recognition unit 506 is used to perform object recognition on the sample image in the second recognition network to obtain a second recognition result corresponding to the target object. The second recognition result is used to indicate the position distribution of the second candidate position set obtained after recognizing the target object. The network structure size of the first recognition network is smaller than the network structure size of the second recognition network.
[0122] The second acquisition unit 508 is used to acquire the first position distribution loss between the first recognition result and the object position label, and the second position distribution loss between the first recognition result and the second recognition result;
[0123] Training unit 510 is used to obtain a target recognition network that has reached the convergence condition after training the first recognition network based on the first position distribution loss and the second position distribution loss, wherein the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold.
[0124] Optionally, in this embodiment, the implementation of each of the above-mentioned unit modules can be referred to the above-mentioned method embodiments, which will not be repeated here.
[0125] According to another aspect of the present invention, an electronic device for implementing the above-described object recognition network training method is also provided. This electronic device may be... Figure 6 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 6 As shown, the electronic device includes a memory 602 and a processor 604. The memory 602 stores a computer program, and the processor 604 is configured to execute the steps in any of the above method embodiments via the computer program.
[0126] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0127] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0128] S1, obtain a set of sample images and a set of object location labels corresponding to the set of sample images, wherein the sample images in the set of sample images contain the target objects to be identified, and the object location labels in the set of object location labels are used to indicate the location labels corresponding to the positions of the target objects in the sample images;
[0129] S2, In the first recognition network, object recognition is performed on the sample image to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the position distribution of the first candidate position set obtained after the target object is recognized;
[0130] S3, In the second recognition network, object recognition is performed on the sample image to obtain a second recognition result corresponding to the target object. The second recognition result is used to indicate the position distribution of the second candidate position set obtained after recognizing the target object. The network structure size of the first recognition network is smaller than that of the second recognition network.
[0131] S4, obtain the first position distribution loss between the first recognition result and the object position label, and the second position distribution loss between the first recognition result and the second recognition result;
[0132] S5. Based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached the convergence condition after training the first recognition network is obtained, wherein the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold.
[0133] Alternatively, as those skilled in the art will understand, Figure 6 The structure shown is for illustrative purposes only. Electronic devices can also be in-vehicle terminals, smartphones (such as Android phones, iOS phones, etc.), tablets, handheld computers, mobile internet devices (MIDs), PADs, and other terminal devices. Figure 6 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 6 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 6 The different configurations shown.
[0134] The memory 602 can be used to store software programs and modules, such as the program instructions / modules corresponding to the object recognition network training method and apparatus in this embodiment of the invention. The processor 604 executes various functional applications and data processing by running the software programs and modules stored in the memory 602, thereby implementing the above-mentioned object recognition network training method. The memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 602 may further include memory remotely located relative to the processor 604, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 602 may be used, but is not limited to, to store information such as various elements in the viewing angle image and training information of the object recognition network. As an example, such as Figure 6As shown, the memory 602 may include, but is not limited to, the first acquisition unit 502, the first recognition unit 504, the second recognition unit 506, the second acquisition unit 508, and the training unit 510 in the training device of the object recognition network. Furthermore, it may include, but is not limited to, other module units in the training device of the object recognition network, which will not be elaborated upon in this example.
[0135] Optionally, the transmission device 606 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 606 includes a Network Interface Controller (NIC), which can be connected to other network devices and routers via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 606 is a radio frequency (RF) module, used for wireless communication with the Internet.
[0136] In addition, the above-mentioned electronic device also includes a display 608 and a connection bus 610 for connecting the various module components in the above-mentioned electronic device.
[0137] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.
[0138] According to one aspect of this application, a computer program product is provided, comprising a computer program / instructions containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit, it performs various functions provided in embodiments of this application.
[0139] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0140] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of a computer device reads computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the training method of the object recognition network described above.
[0141] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0142] S1, obtain a set of sample images and a set of object location labels corresponding to the set of sample images, wherein the sample images in the set of sample images contain the target objects to be identified, and the object location labels in the set of object location labels are used to indicate the location labels corresponding to the positions of the target objects in the sample images;
[0143] S2, In the first recognition network, object recognition is performed on the sample image to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate the position distribution of the first candidate position set obtained after the target object is recognized;
[0144] S3, In the second recognition network, object recognition is performed on the sample image to obtain a second recognition result corresponding to the target object. The second recognition result is used to indicate the position distribution of the second candidate position set obtained after recognizing the target object. The network structure size of the first recognition network is smaller than that of the second recognition network.
[0145] S4, obtain the first position distribution loss between the first recognition result and the object position label, and the second position distribution loss between the first recognition result and the second recognition result;
[0146] S5. Based on the first position distribution loss and the second position distribution loss, a target recognition network that has reached the convergence condition after training the first recognition network is obtained, wherein the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold.
[0147] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0148] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0149] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.
[0151] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0152] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0153] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for training an object recognition network, the method comprising: The method comprises the following steps: obtaining a sample image set and an object position label set corresponding to the sample image set, wherein a target object to be recognized is contained in a sample image in the sample image set, and an object position label in the object position label set is used to indicate a position label corresponding to a position of the target object in the sample image; performing object recognition on the sample image in a first recognition network to obtain a first recognition result corresponding to the target object, wherein the first recognition result is used to indicate a position distribution of a first candidate position set obtained after recognizing the target object; performing object recognition on the sample image in a second recognition network to obtain a second recognition result corresponding to the target object, wherein the second recognition result is used to indicate a position distribution of a second candidate position set obtained after recognizing the target object; the network structure size of the first recognition network is smaller than the network structure size of the second recognition network; obtaining a first position distribution loss between the first recognition result and the object position label, and a second position distribution loss between the first recognition result and the second recognition result; obtaining a target recognition network that reaches a convergence condition after training the first recognition network according to the first position distribution loss and the second position distribution loss, wherein the similarity between the recognition result output by the target recognition network and the recognition result output by the second recognition network is less than a threshold value; The method comprises the following steps:
2. The method of claim 1, wherein, obtaining a target regression loss function and a target classification loss function; calculating a regression loss value between a first numerical distribution sequence corresponding to the first recognition result and an object position coordinate value sequence included in the object position label by using the target regression loss function, and calculating a classification loss value between the first numerical distribution sequence and the object position coordinate value sequence by using the classification loss function; and obtaining a weighted sum value of the regression loss value and the classification loss value as the first position distribution loss. The method comprises the following steps: obtaining the first numerical distribution sequence corresponding to the first recognition result, and a second numerical distribution sequence corresponding to the second recognition result; obtaining the object position coordinate value sequence included in the object position label; determining the first position distribution loss by using the first numerical distribution sequence and the object position coordinate value sequence; 3. The method of claim 2, wherein, determining the second position distribution loss by using the first numerical distribution sequence and the second numerical distribution sequence. The method comprises the following steps: traversing the first numerical distribution sequence and the second numerical distribution sequence, and performing the following steps until the traversal ends: obtain an i-th first distribution value of the first numerical distribution sequence and an i-th second distribution value of the second numerical distribution sequence, where i is a positive integer, the i-th first distribution value corresponds to the i-th second distribution value; input the i-th second distribution value into a temperature coefficient input normalized exponential function to obtain an i-th reference second distribution value, where the temperature coefficient is used to indicate a simulation compactness degree based on the second numerical distribution sequence; obtain a logarithm value of a ratio of the i-th first distribution value and the i-th reference second distribution value as an i-th numerical weight coefficient; obtain a ratio of the i-th numerical weight coefficient and the i-th first distribution value as an i-th loss value; obtain a sum of N loss values as the second position distribution loss in a case that the first numerical distribution sequence and the second numerical distribution sequence are traversed to an end, where N is a number of loss values obtained in a traversal process.
4. The method of claim 1, wherein, the object recognition on the sample image in the first recognition network to obtain a first recognition result corresponding to the target object and the object recognition on the sample image in the second recognition network to obtain a second recognition result corresponding to the target object include: performing object recognition on the sample image by using a first recognition sub-network to obtain a first object position coordinate value sequence; inputting the first object position coordinate value sequence into a mapping sub-network to obtain a first numerical distribution sequence; performing object recognition on the sample image by using a second recognition sub-network to obtain a second object position coordinate value sequence; inputting the second object position coordinate value sequence into the mapping sub-network to obtain a second numerical distribution sequence, where a network parameter size of the first recognition sub-network is smaller than the second recognition sub-network.
5. The method according to any one of claims 1 to 4, characterized in that, the target recognition network that the first recognition network is trained to reach a convergence condition according to the first position distribution loss and the second position distribution loss includes: obtaining a current joint coefficient, where the current joint coefficient is a weight coefficient determined according to a current number of training times that have been performed, and the current joint coefficient is less than 1; determining a current reference joint coefficient according to a difference between a numerical 1 and the current joint coefficient; taking a product of the current reference joint coefficient and the first position distribution loss and a sum of a product of the current joint coefficient and the second position distribution loss as a joint loss value; in a case that the joint loss value is less than or equal to a target loss threshold value, taking a current first recognition network as the target recognition network that reaches the convergence condition; in a case that the joint loss value is greater than the target loss threshold value, continuing network training of the current first recognition network.
6. The method of claim 5, wherein, the obtaining of the current joint coefficient includes: in a case that the number of training times that have been performed is less than or equal to N1, determining the current joint coefficient as a first joint coefficient, where the first joint coefficient is a constant greater than or equal to 0, and N1 is a positive integer; In a case that the number of the training times is greater than N1 and less than N2 or equal to N2, the current joint coefficient is determined as a second joint coefficient, wherein the second joint coefficient increases with the number of the training times, the second joint coefficient is greater than the first joint coefficient, and N2 is a positive integer greater than N1. In a case that the number of the training times is greater than N2 and less than N3 or equal to N3, the current joint coefficient is determined as a third joint coefficient, wherein the third joint coefficient is a constant greater than the second joint coefficient, and N3 is a positive integer greater than N2.
7. A training device for an object recognition network, characterized in that, The method comprises: a first obtaining unit, configured to obtain a sample image set and an object position label set corresponding to the sample image set, wherein a sample image in the sample image set contains a target object to be recognized, and an object position label in the object position label set is used to indicate a position label corresponding to a position of the target object in the sample image; a first identifying unit, configured to perform object recognition on the sample image in a first identification network to obtain a first identification result corresponding to the target object, wherein the first identification result is used to indicate a position distribution of a first candidate position set obtained after the target object is recognized; a second identifying unit, configured to perform object recognition on the sample image in a second identification network to obtain a second identification result corresponding to the target object, wherein the second identification result is used to indicate a position distribution of a second candidate position set obtained after the target object is recognized; and a network structure size of the first identification network is less than a network structure size of the second identification network; a second obtaining unit, configured to obtain a first position distribution loss between the first identification result and the object position label, and a second position distribution loss between the first identification result and the second identification result; a training unit, configured to obtain a target identification network after the first identification network is trained according to the first position distribution loss and the second position distribution loss, wherein a similarity between an identification result output by the target identification network and an identification result output by the second identification network is less than a threshold value; the second obtaining unit is configured to obtain a target regression loss function and a target classification loss function; calculate a regression loss value between a first numerical value distribution sequence corresponding to the first identification result and an object position coordinate value sequence included in the object position label by using the target regression loss function, and calculate a classification loss value between the first numerical value distribution sequence and the object position coordinate value sequence by using the classification loss function; and obtain a weighted sum value of the regression loss value and the classification loss value as the first position distribution loss.
8. A computer readable storage medium, characterized in that, The computer-readable storage medium comprises a stored program, wherein the program performs the method described in any one of claims 1 to 6 when executed.
9. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions, when executed by a processor, implement the steps of the method described in any one of claims 1 to 6.
10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 6 by using the computer program.
Citation Information
Patent Citations
Object recognition method and device, storage medium and electronic equipment
CN112132231A