Ship Management Method Based on Neural Network
By using neural network model to fusion processing of video frame images in the ship management system, the problem of insufficient accuracy of behavior analysis in neural networks in ship scenes is solved, and higher behavior analysis accuracy and safety risk detection capabilities are achieved.
Patent Information
- Application Number
- CN202410829390.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-06-25
AI Technical Summary
In the scenario of high safety requirements for ships, the accuracy of behavioral analysis needs to be further improved.
Using a neural network-based ship management method, by acquiring and cropping video frame images, the neural network model is triggered to process images in a feature fusion manner, and fusing feature vectors of different frames to improve the accuracy of behavioral analysis.
It improves the accuracy and robustness of neural networks to analyze ship-related personnel behaviors, and enhances the ability to detect potential safety risks.
Smart Images

Figure CN118736489B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a ship management method based on neural network. Background Art
[0002] During the voyage of a ship, ensuring the behavior of personnel / seafarers is the key to ensuring navigation safety. With the development of information technology, artificial intelligence (AI) is increasingly used in the field of ships, especially in monitoring and analyzing the behavior of ship personnel / seafarers. By processing and intelligently analyzing videos, artificial intelligence technology can help ship-related personnel promptly detect and deal with behaviors that may affect navigation safety, thereby improving the overall safety level of the ship.
[0003] Specifically, it includes the following technologies:
[0004] 1. Video surveillance system: Modern ships are generally equipped with high-definition cameras to record the situation inside the ship in real time. Through the video surveillance system, ship managers can view the real-time images inside the ship at any time to ensure the transparency of ship operations.
[0005] 2. Behavior analysis technology: Using machine learning and deep learning technologies, the behavior of people / seafarers in the video footage can be analyzed to identify behaviors that do not comply with regulations, such as fatigue driving, leaving the work station without authorization, not wearing a mask, not wearing a helmet, falling, smoking, etc.
[0006] 3. Abnormal behavior alarm system: When the system detects abnormal behavior, it will immediately send an alarm to the ship management personnel so that timely measures can be taken to prevent potential safety risks.
[0007] The above technologies can improve the efficiency of safety management. For example, artificial intelligence technology can monitor the behavior of ship personnel / seafarers in real time, discover problems in a timely manner, and reduce safety accidents caused by human factors. It can also enhance the objectivity of safety management. For example, through objective behavioral analysis of videos, it can reduce human bias and misjudgment, improve the accuracy of safety management, and promote continuous improvement of safety management. For example, artificial intelligence technology can continuously learn and optimize, and through analysis of new data, it can continuously improve safety management measures to adapt to the ever-changing ship operating environment.
[0008] However, the accuracy of behavioral analysis of neural networks in high-safety demand scenarios such as ships needs to be further improved. Summary of the invention
[0009] The embodiment of the present application provides a ship management method based on a neural network to improve the accuracy of the neural network in analyzing the behavior of personnel related to the ship.
[0010] In order to achieve the above objectives, this application adopts the following technical solutions:
[0011] In a first aspect, a ship management method based on a neural network is provided, the method being applied to an electronic device deployed on a ship, the method comprising: the electronic device acquiring L frame images collected for a monitored area, where L is an integer greater than 1, the electronic device performing cropping processing on the L frame images to obtain cropped L frame images; the electronic device triggering a neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of a monitored object in the cropped L frame images; wherein feature fusion means: the electronic device performing convolution on the cropped L frame images through the neural network model to obtain a feature vector of each cropped image in the cropped L frame images, obtaining a total of L feature vectors; the electronic device fusing the L feature vectors through the neural network model to obtain a vector module; the electronic device analyzing the vector module through the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0012] Optionally, L is an integer multiple of 3, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the jth feature vector, the 2jth feature vector and the 3jth feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, j traverses from 1 to L / 3, and a total of L / 3 vector modules are obtained. 3. The method according to claim 2 is characterized in that the jth feature vector includes P feature vector components Uj, and the P feature vector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), the 3jth eigenvector includes P eigenvector components U3j, and the P eigenvector components U3j are expressed as (U3j1, U3j2, …, U3j P ), P is an integer greater than 1;
[0013] The electronic device fuses the jth feature vector, the 2jth feature vector and the 3jth feature vector of the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, including: the electronic device uses Uj1, U2j1 and U3j1 as the first group of feature vector components j, and uses Uj2, U2j2 and U3j2 as the second group of feature vector components j through the fusion layer of the neural network model, and so on, until Uj P 、U2jP and U3j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device uses the fusion layer of the neural network model to take each group of feature vector components j in the P groups of feature vector components j as a pixel point j to obtain the j-th feature image containing P pixel points j, j traverses 1 to L / 3, and a total of L / 3 feature images are obtained; the electronic device convolves the L / 3 feature images through the convolution layer of the neural network model to obtain L / 3 feature feature vectors; the electronic device analyzes the L / 3 feature feature vectors through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0014] Optionally, L is an integer multiple of 2, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the j-th feature vector and the 2j-th feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the j-th vector module, j traverses 1 to L / 2, and a total of L / 2 vector modules are obtained.
[0015] Optionally, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, ..., Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), P is an integer greater than 1; the electronic device uses the fusion layer of the neural network model to obtain the jth feature vector and the 2jth feature vector in the L feature vectors to obtain the jth vector module, including: the electronic device uses the fusion layer of the neural network model to use Uj1 and U2j1 as the first group of feature vector components j, and uses Uj2 and U2j2 as the second group of feature vector components j, and so on, until Uj P and U2j PAs the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device analyzes the L / 2 vector modules through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0016] Optionally, the electronic device obtains L frames of images collected for the monitored area, including: the electronic device obtains the L frames of images from the video collected from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2.
[0017] Optionally, the electronic device obtains the L frame image from the video captured from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2, including: the electronic device determines that the video captured from the monitored area contains a total of W frame images, W is less than L and is not an integer multiple of 3 / 2; the electronic device uniformly inserts LW frame images into the W frame images to obtain the L frame image.
[0018] Optionally, the electronic device crops the L frame image to obtain a cropped L frame image, including: if the monitored area is located on the left side of the image, the electronic device crops the right part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located on the right side of the image, the electronic device crops the left part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located in the center of the image, the electronic device crops the outer part of each frame image in the L frame image to obtain the cropped L frame image.
[0019] Optionally, the behavior of the monitored object includes at least one of the following: not wearing a mask, not wearing a helmet, falling, and smoking.
[0020] In a second aspect, a neural network-based ship management device is provided, which is applied to an electronic device deployed on a ship, and the device is configured as follows: the electronic device obtains L frame images collected for the monitored area, N is an integer greater than 1, and the electronic device crops the cropped L frame images to obtain the cropped L frame images; the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images; wherein feature fusion means: the electronic device convolves the cropped L frame images through the neural network model to obtain a feature vector of each cropped image in the cropped L frame images, and obtains a total of L feature vectors; the electronic device fuses the L feature vectors through the neural network model to obtain a vector module; the electronic device analyzes the vector module through the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0021] Optionally, L is an integer multiple of 3, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the j-th feature vector, the 2j-th feature vector and the 3j-th feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the j-th vector module, j traverses 1 to L / 3, and a total of L / 3 vector modules are obtained.
[0022] Optionally, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), the 3jth eigenvector includes P eigenvector components U3j, and the P eigenvector components U3j are expressed as (U3j1, U3j2, …, U3j P ), P is an integer greater than 1; the electronic device fuses the jth feature vector, the 2jth feature vector and the 3jth feature vector in the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, including: the electronic device uses Uj1, U2j1 and U3j1 as the first group of feature vector components j, and uses Uj2, U2j2 and U3j2 as the second group of feature vector components j through the fusion layer of the neural network model, and so on, until Uj P 、U2j P and U3j PAs the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device uses the fusion layer of the neural network model to take each group of feature vector components j in the P groups of feature vector components j as a pixel point j to obtain the j-th feature image containing P pixel points j, j traverses 1 to L / 3, and a total of L / 3 feature images are obtained; the electronic device convolves the L / 3 feature images through the convolution layer of the neural network model to obtain L / 3 feature feature vectors; the electronic device analyzes the L / 3 feature feature vectors through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0023] Optionally, L is an integer multiple of 2, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the j-th feature vector and the 2j-th feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the j-th vector module, j traverses 1 to L / 2, and a total of L / 2 vector modules are obtained.
[0024] Optionally, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, ..., Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), P is an integer greater than 1; the electronic device uses the fusion layer of the neural network model to obtain the jth feature vector and the 2jth feature vector in the L feature vectors to obtain the jth vector module, including: the electronic device uses the fusion layer of the neural network model to use Uj1 and U2j1 as the first group of feature vector components j, and uses Uj2 and U2j2 as the second group of feature vector components j, and so on, until Uj P and U2j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device analyzes the L / 2 vector modules through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0025] Optionally, the electronic device obtains L frames of images collected for the monitored area, including: the electronic device obtains the L frames of images from the video collected from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2.
[0026] Optionally, the electronic device obtains the L frame image from the video captured from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2, including: the electronic device determines that the video captured from the monitored area contains a total of W frame images, W is less than L and is not an integer multiple of 3 / 2; the electronic device uniformly inserts LW frame images into the W frame images to obtain the L frame image.
[0027] Optionally, the electronic device crops the L frame image to obtain a cropped L frame image, including: if the monitored area is located on the left side of the image, the electronic device crops the right part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located on the right side of the image, the electronic device crops the left part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located at the center of the image, the electronic device crops the outer part of each frame image in the L frame image to obtain the cropped L frame image.
[0028] Optionally, the behavior of the monitored object includes at least one of the following: not wearing a mask, not wearing a helmet, falling, and smoking.
[0029] In summary, the above method and device have the following technical effects:
[0030] After the electronic device obtains L frames of images collected for the monitored area and crops them, the electronic device can trigger the neural network model to process the cropped L frames of images in a feature fusion manner. For example, the neural network model can extract features of the L frame images, such as L feature vectors. The neural network model can fuse the L feature vectors to obtain a vector module, so that the behavior changes of the same object in images of different frames can be fused together for analysis, thereby improving the accuracy or robustness of the neural network's analysis of ship-related personnel behavior. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] Figure 1 A schematic diagram of a process flow of an AI-based ship management and monitoring method provided in an embodiment of the present application;
[0032] Figure 2 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In the embodiment of the present invention, "indication" may include direct indication and indirect indication, and may also include explicit indication and implicit indication. The information indicated by a certain information is called information to be indicated. In the specific implementation process, there are many ways to indicate the information to be indicated, such as but not limited to, the information to be indicated can be directly indicated, such as the information to be indicated itself or the index of the information to be indicated. The information to be indicated can also be indirectly indicated by indicating other information, wherein there is an association relationship between the other information and the information to be indicated. It is also possible to indicate only a part of the information to be indicated, while the other parts of the information to be indicated are known or agreed in advance. For example, the indication of specific information can also be achieved by means of the arrangement order of each information agreed in advance (for example, specified by the protocol), thereby reducing the indication overhead to a certain extent. At the same time, the common parts of each information can also be identified and indicated uniformly to reduce the indication overhead caused by indicating the same information separately.
[0034] In addition, the specific indication method may also be various existing indication methods, such as but not limited to the above-mentioned indication methods and various combinations thereof. The specific details of the various indication methods can refer to the prior art and will not be repeated herein. As can be seen from the above, for example, when it is necessary to indicate multiple information of the same type, different indication methods may be used for different information. In the specific implementation process, the desired indication method can be selected according to specific needs. The embodiment of the present invention does not limit the selected indication method. In this way, the indication method involved in the embodiment of the present invention should be understood to cover various methods that can enable the party to be indicated to obtain the information to be indicated.
[0035] It should be understood that the information to be indicated can be sent as a whole or divided into multiple sub-information and sent separately, and the sending period and / or sending time of these sub-information can be the same or different. The specific sending method is not limited in the embodiment of the present invention. Among them, the sending period and / or sending time of these sub-information can be pre-defined, for example, pre-defined according to a protocol, or can be configured by the sending end device by sending configuration information to the receiving end device.
[0036] "Pre-definition" or "pre-configuration" can be implemented by pre-saving corresponding codes, tables or other methods that can be used to indicate relevant information in the device, and the embodiments of the present invention do not limit the specific implementation method. Among them, "saving" can mean saving in one or more memories. The one or more memories can be set separately or integrated in an encoder or decoder, a processor, or an electronic device. The one or more memories can also be partially set separately and partially integrated in a decoder, a processor, or an electronic device. The type of memory can be any form of storage medium, which is not limited by the embodiments of the present invention.
[0037] The "protocol" involved in the embodiments of the present invention may refer to a protocol family in the communication field, a standard protocol with a similar protocol family frame structure, or a related protocol in a reliable access method system for future Internet of Things devices, and the embodiments of the present invention do not specifically limit this.
[0038] In the embodiments of the present invention, descriptions such as "when...", "in the case of...", "if" and "if" all mean that the device will perform corresponding processing under certain objective circumstances, but do not limit the time, nor do they require the device to perform judgment actions when implementing, nor do they mean the existence of other limitations.
[0039] In the description of the embodiments of the present invention, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B can represent A or B; "and / or" in the embodiments of the present invention is only a kind of association relationship describing the associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In addition, in the description of the embodiments of the present invention, unless otherwise specified, "multiple" refers to two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple. In addition, in order to facilitate the clear description of the technical solution of the embodiments of the present invention, in the embodiments of the present invention, the words "first" and "second" are used to distinguish the same or similar items with basically the same functions and effects. Those skilled in the art will appreciate that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit the difference. At the same time, in the embodiments of the present invention, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0040] The network architecture and business scenarios described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. A person of ordinary skill in the art can appreciate that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0041] The technical solution in this application will be described below in conjunction with the accompanying drawings.
[0042] See also Figure 1 , the embodiment of the present application provides a ship management method based on a neural network. The method is executed by an electronic device, and the process of the method includes:
[0043] S101, the electronic device obtains L frames of images collected for a monitored area, where L is an integer greater than 1.
[0044] The electronic equipment is deployed on the ship, which can be understood as a terminal. The terminal can be a terminal with communication function, or a chip or chip system set in the terminal. The terminal device can also be called user equipment (UE), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device. The terminal device in the embodiment of the present application can be a mobile phone, a tablet computer (Pad), a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a wireless terminal in self driving, a wireless terminal in remote medical, a wireless terminal in smart grid, a wireless terminal in transportation safety, a wireless terminal in smart city, a wireless terminal in smart home, a vehicle-mounted terminal, an RSU with terminal function, etc.
[0045] L is an integer multiple of 3 or L is an integer multiple of 2. On this basis, the electronic device can obtain L frame images from the video collected in the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2. For example, in a specific case, the electronic device determines that the video collected in the monitored area contains a total of W frame images, W is less than L and is not an integer multiple of 3 / 2; the electronic device uniformly inserts LW frame images into the W frame images to obtain L frame images, that is, at the position where the frame needs to be inserted, the image of the adjacent frame is copied and inserted into the position to make up the L frame image. Alternatively, if the video collected in the monitored area contains a total of W frame images, and W is greater than L, the electronic device uniformly extracts L frame images from the W frame images.
[0046] S102, the electronic device performs cropping processing on the L frame image to obtain a cropped L frame image.
[0047] For example, if the monitored area is located on the left side of the image, the electronic device may crop the right part of each frame image in the L frame image to obtain a cropped L frame image, i.e., only retaining the monitored area; if the monitored area is located on the right side of the image, the electronic device may crop the left part of each frame image in the L frame image to obtain a cropped L frame image, i.e., only retaining the monitored area; if the monitored area is located in the center of the image, the electronic device may crop the outer part of each frame image in the L frame image to obtain a cropped L frame image, i.e., only retaining the monitored area.
[0048] S103, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images.
[0049] First, the electronic device can convolve the cropped L frame images through the neural network model to obtain the feature vector of each cropped image in the cropped L frame images, and obtain L feature vectors in total. For example, the electronic device can convolve the cropped L frame images separately through the convolution layer of the neural network model (recorded as convolution layer #1). Convolution layer #1 can include one or more convolution layers, and one or more pooling layers, such as the processing architecture is: convolution layer → pooling layer → convolution layer → pooling layer. The N feature vectors can continue to be passed to the fusion layer of the neural network model.
[0050] Secondly, the electronic device fuses L feature vectors through a neural network model to obtain a vector module.
[0051] In a possible embodiment: for the case where L is an integer multiple of 3.
[0052] The electronic device can fuse the jth feature vector, the 2jth feature vector, and the 3jth feature vector in the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, and j traverses from 1 to L / 3, and a total of L / 3 vector modules are obtained. For example, the jth feature vector may include P feature vector components Uj, and the P feature vector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1,U2j2,…,U2j P ), the 3jth eigenvector may include P eigenvector components U3j, and the P eigenvector components U3j are expressed as (U3j1, U3j2, …, U3j P), P is an integer greater than 1. On this basis, the electronic device uses the fusion layer of the neural network model to use Uj1, U2j1 and U3j1 as the first group of feature vector components j, and use Uj2, U2j2 and U3j2 as the second group of feature vector components j, and so on, until Uj P 、U2j P and U3j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module, that is, the feature vectors of every 3 frames of images are fused into one result.
[0053] In another possible embodiment: for the case where L is an integer multiple of 2.
[0054] The electronic device can obtain the jth vector module by combining the jth feature vector and the 2jth feature vector in the L feature vectors through the fusion layer of the neural network model, and j traverses from 1 to L / 2 to obtain a total of L / 2 vector modules. For example, the jth feature vector may include P feature vector components Uj, and the P feature vector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector may include P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), P is an integer greater than 1; on this basis, the electronic device can use the fusion layer of the neural network model to use Uj1 and U2j1 as the first group of feature vector components j, and use Uj2 and U2j2 as the second group of feature vector components j, and so on, until Uj P and U2j P As the Pth group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the jth vector modules, that is, the feature vectors of every two frames of images are fused into one result. .
[0055] Finally, the electronic device analyzes the vector module through a neural network model to determine the behavior of the monitored object in the cropped L-frame image, such as if the behavior of the monitored object includes at least one of the following: not wearing a mask, not wearing a helmet, falling, smoking, etc.
[0056] In a possible embodiment of the above:
[0057] The electronic device can continue to use the fusion layer of the neural network model to treat each set of feature vector components j in the P sets of feature vector components j as a pixel point j, and obtain the jth feature image containing P pixel points j. j traverses 1 to L / 3, and obtains a total of L / 3 feature images. For example, the three feature vector components j in each set of feature vector components j correspond to the RGB values of a pixel point j, and P pixel points j are obtained. Then the P pixel points j are arranged in order into a matrix, that is, the jth feature image containing P pixel points j is obtained, and the size of the matrix is the image resolution. The electronic device then convolves the L / 3 feature images through the convolution layer of the neural network model (different from the above-mentioned convolution layer #1, recorded as convolution layer #2), and obtains L / 3 feature vectors. It can be understood that convolution layer #2 usually does not include a pooling layer, that is, only one more convolution is performed on the feature image. The advantage of this is that deeper feature extraction is performed through convolution, thereby improving the robustness of subsequent fully connected layer processing. In this way, the electronic device can analyze the L / 3 characteristic feature vectors through the fully connected layer of the neural network model to determine the behavior of the object.
[0058] In another possible embodiment of the above:
[0059] The electronic device can also directly analyze L / 2 vector modules through the fully connected layer of the neural network model to determine whether the object in the video stream has any illegal / abnormal behavior. Alternatively, in this case, it can also be converted into a feature image, and then processed similarly to a possible embodiment. For example, two feature vector components j in each group of feature vector components j correspond one-to-one as two values in the RGB value of a pixel j, and the other value in the RGB value is filled with 0, and P pixel points j can also be obtained. Then, the P pixel points j are arranged in order into a matrix, that is, the jth feature image containing P pixel points j is obtained, and the size of the matrix is the image resolution.
[0060] It can be understood that through feature fusion, the analysis of the neural network can take into account the behavioral changes of objects (such as people) in different images, thereby improving the robustness of the neural network model analysis.
[0061] In summary, after the electronic device obtains L frames of images collected for the monitored area and crops them, the electronic device can trigger the neural network model to process the cropped L frames of images in a feature fusion manner. For example, the neural network model can extract features of the L frame images, such as L feature vectors. The neural network model can fuse the L feature vectors to obtain a vector module, so that the behavior changes of the same object in images of different frames can be fused together for analysis, thereby improving the accuracy or robustness of the neural network's analysis of ship-related personnel behavior.
[0062] Combination of the above Figure 1 The neural network-based ship management method provided in an embodiment of the present application is described in detail. The following introduces a neural network-based ship management device applying the method, which is applied to an electronic device deployed on a ship, and the device is configured as follows: the electronic device obtains L frame images collected for the monitored area, N is an integer greater than 1, and the electronic device performs cropping processing on the cropped L frame image to obtain the cropped L frame image; the electronic device triggers the neural network model to process the cropped L frame image in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame image; wherein feature fusion means: the electronic device convolves the cropped L frame image through the neural network model to obtain a feature vector of each cropped image in the cropped L frame image, and obtains a total of L feature vectors; the electronic device fuses the L feature vectors through the neural network model to obtain a vector module; the electronic device analyzes the vector module through the neural network model to determine the behavior of the monitored object in the cropped L frame image.
[0063] Optionally, L is an integer multiple of 3, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the j-th feature vector, the 2j-th feature vector and the 3j-th feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the j-th vector module, j traverses 1 to L / 3, and a total of L / 3 vector modules are obtained.
[0064] Optionally, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), the 3jth eigenvector includes P eigenvector components U3j, and the P eigenvector components U3j are expressed as (U3j1, U3j2, …, U3j P ), P is an integer greater than 1; the electronic device fuses the jth feature vector, the 2jth feature vector and the 3jth feature vector in the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, including: the electronic device uses Uj1, U2j1 and U3j1 as the first group of feature vector components j, and uses Uj2, U2j2 and U3j2 as the second group of feature vector components j through the fusion layer of the neural network model, and so on, until Uj P 、U2j Pand U3j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device uses the fusion layer of the neural network model to take each group of feature vector components j in the P groups of feature vector components j as a pixel point j to obtain the j-th feature image containing P pixel points j, j traverses 1 to L / 3, and a total of L / 3 feature images are obtained; the electronic device convolves the L / 3 feature images through the convolution layer of the neural network model to obtain L / 3 feature feature vectors; the electronic device analyzes the L / 3 feature feature vectors through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0065] Optionally, L is an integer multiple of 2, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: the electronic device fuses the j-th feature vector and the 2j-th feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the j-th vector module, j traverses 1 to L / 2, and a total of L / 2 vector modules are obtained.
[0066] Optionally, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, ..., Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), P is an integer greater than 1; the electronic device uses the fusion layer of the neural network model to obtain the jth feature vector and the 2jth feature vector in the L feature vectors to obtain the jth vector module, including: the electronic device uses the fusion layer of the neural network model to use Uj1 and U2j1 as the first group of feature vector components j, and uses Uj2 and U2j2 as the second group of feature vector components j, and so on, until Uj P and U2j PAs the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; accordingly, the electronic device triggers the neural network model to process the cropped L frame images in a feature fusion manner to determine the behavior of the monitored object in the cropped L frame images, including: the electronic device analyzes the L / 2 vector modules through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
[0067] Optionally, the electronic device obtains L frames of images collected for the monitored area, including: the electronic device obtains the L frames of images from the video collected from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2.
[0068] Optionally, the electronic device obtains the L frame image from the video captured from the monitored area with L being an integer multiple of 3 or L being an integer multiple of 2, including: the electronic device determines that the video captured from the monitored area contains a total of W frame images, W is less than L and is not an integer multiple of 3 / 2; the electronic device uniformly inserts LW frame images into the W frame images to obtain the L frame image.
[0069] Optionally, the electronic device crops the L frame image to obtain a cropped L frame image, including: if the monitored area is located on the left side of the image, the electronic device crops the right part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located on the right side of the image, the electronic device crops the left part of each frame image in the L frame image to obtain the cropped L frame image; if the monitored area is located in the center of the image, the electronic device crops the outer part of each frame image in the L frame image to obtain the cropped L frame image.
[0070] Optionally, the behavior of the monitored object includes at least one of the following: not wearing a mask, not wearing a helmet, falling, and smoking.
[0071] Combine the following Figure 2 The components of the electronic device 500 are described in detail:
[0072] The processor 501 is the control center of the electronic device 500, and may be a processor or a general term for multiple processing elements. For example, the processor 501 is one or more central processing units (CPUs), or may be application specific integrated circuits (ASICs), or may be one or more integrated circuits configured to implement the embodiments of the present application, such as one or more microprocessors (digital signal processors, DSPs), or one or more field programmable gate arrays (FPGAs).
[0073] Optionally, the processor 501 can execute various functions of the electronic device 500 by running or executing the software program stored in the memory 502 and calling the data stored in the memory 502, as described above. Figure 1 Functionality in the method shown.
[0074] In a specific implementation, as an embodiment, the processor 501 may include one or more CPUs, such as Figure 2 CPU0 and CPU1 are shown in FIG.
[0075] In a specific implementation, as an embodiment, the electronic device 500 may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor here may refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0076] Among them, the memory 502 is used to store the software program for executing the solution of the present application, and the execution is controlled by the processor 501. The specific implementation method can refer to the above method embodiment and will not be repeated here.
[0077] Optionally, the memory 502 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or
[0078] Other types of dynamic storage devices that can store information and instructions may also be electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical disc, laser disc, optical disc, digital versatile disc, Blu-ray disc, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these. The memory 502 may be integrated with the processor 501, or it may exist independently and be stored in the electronic device 500.
[0079] The interface circuit ( Figure 2 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.
[0080] The transceiver 503 is used for communication with other devices. For example, if the multi-beam based positioning device is a terminal, the transceiver 503 can be used for communication with a network device or another terminal.
[0081] Optionally, the transceiver 503 may include a receiver and a transmitter ( Figure 2 The receiver is used to implement a receiving function, and the transmitter is used to implement a sending function.
[0082] Optionally, the transceiver 503 may be integrated with the processor 501, or may exist independently and communicate with the processor 501 through the interface circuit ( Figure 2 (not shown) is coupled to the processor 501, which is not specifically limited in this embodiment of the present application.
[0083] It should be noted that Figure 2 The structure of the electronic device 500 shown in the figure does not constitute a limitation on the device, and the actual electronic device 500 may include more or less components than those shown in the figure, or combine certain components, or arrange the components differently.
[0084] In addition, the technical effects based on the electronic device 500 can refer to the technical effects of the method in the above method embodiment, which will not be repeated here.
[0085] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0086] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0087] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.
[0088] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.
[0089] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0090] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0091] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0092] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0093] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some feature fields can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0094] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0095] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0096] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program codes.
[0097] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A ship management method based on neural network, characterized in that: The method is applied to electronic equipment deployed on a ship, and the method comprises: The electronic device acquires L frames of images collected for the monitored area, where L is an integer greater than 1. The electronic device performs cropping processing on the L frame images to obtain a cropped L frame image; The electronic device triggers the neural network model to process the cropped L-frame image in a feature fusion manner to determine the behavior of the monitored object in the cropped L-frame image; Feature fusion refers to: The electronic device performs convolution on the L cropped image frames through the neural network model to obtain a feature vector of each cropped image in the L cropped image frames, and obtains L feature vectors in total; The electronic device fuses the L feature vectors through the neural network model to obtain a vector module; The electronic device analyzes the vector module through the neural network model to determine the behavior of the monitored object in the cropped L-frame image; L is an integer multiple of 3, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: The electronic device fuses the jth feature vector, the 2jth feature vector, and the 3jth feature vector of the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, and j traverses from 1 to L / 3 to obtain L / 3 vector modules in total; or, L is an integer multiple of 2, and the electronic device fuses the L feature vectors through the fusion layer of the neural network model to obtain a vector module, including: The electronic device fuses the jth feature vector and the 2jth feature vector among the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, and j traverses from 1 to L / 2 to obtain a total of L / 2 vector modules.
2. The method according to claim 1, characterized in that L is an integer multiple of 3, the j-th eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, ..., Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), the 3jth eigenvector includes P eigenvector components U3j, and the P eigenvector components U3j are expressed as (U3j1, U3j2, …, U3j P ), P is an integer greater than 1; The electronic device fuses the jth feature vector, the 2jth feature vector, and the 3jth feature vector of the L feature vectors through the fusion layer of the neural network model to obtain the jth vector module, including: The electronic device uses the fusion layer of the neural network model to take Uj1, U2j1 and U3j1 as the first group of feature vector components j, and takes Uj2, U2j2 and U3j2 as the second group of feature vector components j, and so on, until Uj P 、U2j P and U3j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; Accordingly, the electronic device triggers the neural network model to process the cropped L-frame image in a feature fusion manner to determine the behavior of the monitored object in the cropped L-frame image, including: The electronic device uses the fusion layer of the neural network model to take each set of feature vector components j in the P sets of feature vector components j as a pixel point j, obtains a j-th feature image containing P pixel points j, and traverses j from 1 to L / 3, obtaining a total of L / 3 feature images; The electronic device convolves the L / 3 feature images through the convolution layer of the neural network model to obtain L / 3 feature feature vectors; The electronic device analyzes the L / 3 characteristic feature vectors through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame images.
3. The method according to claim 1, characterized in that L is an integer multiple of 2, and the jth eigenvector includes P eigenvector components Uj, and the P eigenvector components Uj are expressed as (Uj1, Uj2, …, Uj P ), the 2jth eigenvector includes P eigenvector components U2j, and the P eigenvector components U2j are expressed as (U2j1, U2j2, …, U2j P ), P is an integer greater than 1; The electronic device obtains a jth vector module by combining the jth feature vector and the 2jth feature vector of the L feature vectors through the fusion layer of the neural network model, including: The electronic device uses the fusion layer of the neural network model to take Uj1 and U2j1 as the first group of feature vector components j, and takes Uj2 and U2j2 as the second group of feature vector components j, and so on, until Uj P and U2j P As the P-th group of feature vector components j, a total of P groups of feature vector components j are obtained, and the P groups of feature vector components j are the j-th vector module; Accordingly, the electronic device triggers the neural network model to process the cropped L-frame image in a feature fusion manner to determine the behavior of the monitored object in the cropped L-frame image, including: The electronic device analyzes the L / 2 vector modules through the fully connected layer of the neural network model to determine the behavior of the monitored object in the cropped L frame image.
4. The method according to claim 1, characterized in that: The electronic device acquires L frames of images collected for the monitored area, including: The electronic device obtains the L frames of images from the video collected from the monitored area when L is an integer multiple of 3 or L is an integer multiple of 2.
5. The method according to claim 1, characterized in that The electronic device performs cropping processing on the L frame images to obtain a cropped L frame image; If the monitored area is located on the left side of the image, the electronic device cuts off the right side of each frame of the L frames of images to obtain the cut L frames of images; If the monitored area is located on the right side of the image, the electronic device cuts off the left part of each frame of the L frame images to obtain the cut L frame images; If the monitored area is located at the center of the image, the electronic device cuts off the peripheral portion of each frame of the L frame images to obtain the cut L frame images.
6. The method according to claim 1, characterized in that The behavior of the monitored object includes at least one of the following: not wearing a mask, not wearing a helmet, falling, and smoking.
Citation Information
Patent Citations
Behavior detection method and device and electronic equipment
CN117238022A
Football match behavior recognition method and apparatus based on deep learning, and terminal device
WO2020258498A1