A product technical material management method, device, equipment and medium

CN115496057BActive Publication Date: 2026-08-21DEEPAL AUTOMOBILE TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202211262188.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-14
Publication Date
2026-08-21
Estimated Expiration
2042-10-14

AI Technical Summary

Technical Problem

上述现有技术实现了对图书图像文字的识别,但是没有对图书中的每一个区域内容进行归纳,同时也并未将识别结果与初始资料相关联,不利于展示和检索

Benefits of technology

[0044]First, product technical data is acquired. Then, based on the contextual outline of the product technical data, it is divided into multiple data regions, and keywords describing the content of these regions are extracted. Next, the multiple data regions are further divided into multiple text sub-regions and image sub-regions based on their content. The text sub-regions are converted into electronic text data, and the image sub-regions are converted into image data. Finally, a relationship network is constructed based on the keywords, electronic text data, and image data. This relationship network demonstrates the relationships between the product technical data and the keywords, electronic text data, and image data. This invention divides product technical data into regions, extracts keywords from the divided regions, and constructs a relationship network based on the relationship between the keywords and the content of the product technical data. This digitizes the product technical data while improving display and retrieval capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115496057B_ABST
    Figure CN115496057B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of file digitization, and provides a product technical material management method, device, equipment and medium, which comprises the following steps: acquiring product technical material; performing regional division on the product technical material according to the context contour in the product technical material, obtaining a plurality of material regions, and extracting subject words for describing the contents of the plurality of material regions; dividing the plurality of material regions according to the contents in the plurality of material regions, obtaining a plurality of text sub-regions and image sub-regions, converting the plurality of text sub-regions into electronic text data, and converting the plurality of image sub-regions into image data; and constructing a relationship network according to the subject words, the electronic text data and the image data. The application performs regional division on the product technical material, extracts subject words from the divided regions, and constructs a relationship network through the relationship between the subject words and the content of the product technical material, so that the product technical material can be digitized and the display and retrieval capabilities can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of document digitization technology, specifically to a method, apparatus, equipment, and medium for managing product technical data. Background Technology

[0002] With the continuous development of the automotive industry, the increasing electrification and intelligence of automobiles has led to an increase in electronic components in vehicles, forcing OEMs to manage various materials more meticulously, enrich their material information, strengthen material management, and facilitate retrieval and access by technical development personnel within the company.

[0003] Chinese patent CN106250830A discloses a method for structured analysis and processing of digital books. It processes scanned images of books using image processing methods and OCR tools, and represents the books in a structured manner based on their layout, functional, and visual features. Chinese patent CN112905733A discloses a method, system, and apparatus for preserving books based on OCR recognition technology. This method divides text blocks in image files and extracts text features to obtain the recognized text. While these existing technologies achieve text recognition in book images, they do not summarize the content of each region within the book, nor do they correlate the recognition results with the initial data, which is detrimental to display and retrieval. Therefore, how to summarize the content of regions in the initial data and correlate the recognition results with the initial data is a problem that urgently needs to be solved. Summary of the Invention

[0004] In view of the shortcomings of the prior art described above, the purpose of this application is to provide a product technical data management method, apparatus, equipment and medium to solve the problem in the prior art of how to summarize the regional content in the initial data and associate the identification results with the initial data.

[0005] To achieve the above and other related objectives, this application provides a product technical data management method, the method comprising:

[0006] Obtain product technical information;

[0007] Based on the contextual outline in the product technical data, the product technical data is divided into regions to obtain multiple data regions, and keywords used to describe the content of the multiple data regions are extracted.

[0008] The multiple data regions are divided according to their contents to obtain multiple text sub-regions and multiple image sub-regions. The multiple text sub-regions are then converted into electronic text data, and the multiple image sub-regions are converted into image data.

[0009] A relationship network is constructed based on the keywords, electronic text data, and image data. This relationship network is used to demonstrate the relationship between the product technical information and the keywords, electronic text data, and image data.

[0010] In one embodiment of this application, the product technical data includes an electronic component specification sheet, which includes multiple mutually spaced content clusters. Before dividing the product technical data into regions based on the context outline in the product technical data, the method further includes:

[0011] Input the electronic component specifications into the pre-built edge detection model;

[0012] The edge detection model is used to perform convolution and pooling on the multiple content concentration regions to obtain multiple initial contours of different sizes corresponding to different content concentration regions.

[0013] Based on the pre-trained weight parameters, multiple initial contours of different sizes corresponding to each content set region are weighted and fused to obtain context contours corresponding to multiple content set regions.

[0014] In one embodiment of this application, the extraction of subject terms used to describe the content of the plurality of data regions includes:

[0015] If the content in the data area is text, then the data area is input into a pre-built keyword extraction model;

[0016] The topic word extraction model maps each word in the text to a feature vector, and merges the feature vectors of each word to obtain a multi-dimensional feature vector.

[0017] The probability of each word appearing in the multidimensional feature vector is calculated, and the topic words of the content of the multiple data regions are obtained according to the matching relationship between the preset topic words and the frequency of each word.

[0018] In one embodiment of this application, converting the plurality of text sub-regions into electronic text data and converting the plurality of image sub-regions into image data includes:

[0019] The multiple text sub-regions are subjected to text recognition, the recognized text content is converted into an electronic text document, and a document number is assigned to the electronic text document to obtain the electronic text data;

[0020] The image content within the image sub-region is cropped according to the preset image size to obtain the image data.

[0021] In one embodiment of this application, constructing a relationship network based on the keywords, electronic text data, and image data includes:

[0022] Obtain the first mapping relationship between keywords, electronic text data, and image data in the same product's technical documentation;

[0023] Based on the keywords, electronic text data, image data, and the first mapping relationship, construct a relationship network for the same product's technical information;

[0024] If multiple product technical documents contain the same product manufacturing address, then the relationship network of the multiple product technical documents is connected through the product manufacturing location.

[0025] In one embodiment of this application, after constructing the relationship network based on the keywords, electronic text data, and image data, the method further includes:

[0026] Obtain the product model from the product technical data;

[0027] All characters in the electronic text data are segmented to obtain multiple words;

[0028] A second mapping relationship is constructed between the multiple terms, product models, subject terms, and document numbers in the same product technical documents, and the multiple terms, product models, subject terms, and document numbers are stored to obtain a retrieval database.

[0029] In one embodiment of this application, after obtaining the retrieval database, the method further includes:

[0030] Obtain search keywords, which include search topic terms, and / or search product models, and / or search terms. The search priority of search product models is higher than that of search topic terms, and the search priority of search topic terms is higher than that of search terms.

[0031] Based on the search keywords and their priority, information in the search database is retrieved to obtain initial search results, which include product model, and / or subject terms, and / or multiple words, and / or document number.

[0032] Obtain the number of times the product model appears in the initial search results, and / or the number of times the keyword appears, and / or the number of times multiple words appear, and / or the number of times the document number appears;

[0033] The display order of the contents in the initial search results is sorted according to the frequency of occurrence, and the electronic text data corresponding to the initial search results is obtained according to the second mapping relationship.

[0034] In one embodiment of this application, a product technical data management device is also provided, the device comprising:

[0035] The data acquisition module is used to acquire product technical data;

[0036] The keyword acquisition module is used to divide the product technical data into regions based on the context outline in the product technical data, obtain multiple data regions, and extract keyword terms to describe the content of the multiple data regions.

[0037] The electronic text data and image data acquisition module is used to divide the multiple data regions according to the content of the multiple data regions to obtain multiple text sub-regions and multiple image sub-regions, and convert the multiple text sub-regions into electronic text data and the multiple image sub-regions into image data;

[0038] The relationship network construction module is used to construct a relationship network based on the keywords, electronic text data, and image data. The relationship network is used to display the relationship between the product technical information and the keywords, electronic text data, and image data.

[0039] In one embodiment of this application, an electronic device is also provided, the electronic device comprising:

[0040] One or more processors;

[0041] A storage device for storing one or more programs, which, when executed by one or more processors, enable the electronic device to implement the product technical data management method as described above.

[0042] In one embodiment of this application, a computer-readable storage medium is also provided, on which a computer program is stored, which, when executed by a computer's processor, causes the computer to perform the product technical data management method as described above.

[0043] The beneficial effects of this invention are:

[0044] First, product technical data is acquired. Then, based on the contextual outline of the product technical data, it is divided into multiple data regions, and keywords describing the content of these regions are extracted. Next, the multiple data regions are further divided into multiple text sub-regions and image sub-regions based on their content. The text sub-regions are converted into electronic text data, and the image sub-regions are converted into image data. Finally, a relationship network is constructed based on the keywords, electronic text data, and image data. This relationship network demonstrates the relationships between the product technical data and the keywords, electronic text data, and image data. This invention divides product technical data into regions, extracts keywords from the divided regions, and constructs a relationship network based on the relationship between the keywords and the content of the product technical data. This digitizes the product technical data while improving display and retrieval capabilities.

[0045] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0046] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. It is obvious that the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort. In the drawings:

[0047] Figure 1 This is a schematic diagram illustrating the implementation environment of a product technical data management method according to an exemplary embodiment of this application;

[0048] Figure 2 This is a flowchart illustrating a product technical data management method in an exemplary embodiment of this application;

[0049] Figure 3 This is a schematic diagram of a relational network shown in an exemplary embodiment of this application;

[0050] Figure 4 This is a schematic diagram illustrating the interface architecture between the specification digital management system and external systems, as shown in an exemplary embodiment of this application.

[0051] Figure 5 This is a block diagram illustrating a product technical data management device according to an exemplary embodiment of this application;

[0052] Figure 6 A schematic diagram of the structure of a computer system suitable for an electronic device according to an embodiment of this application is shown. Detailed Implementation

[0053] The embodiments of the present invention will be described below with reference to the accompanying drawings and preferred embodiments. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for illustrating the present invention and not for limiting the scope of protection of the present invention.

[0054] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0055] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0056] First, it's important to note that with the continuous development of the automotive industry, the increasing electrification and intelligence of vehicles has led to a surge in electronic components. When developing new products, engineers need to obtain relevant information about the electronic components to be used, such as component specifications. When these components are incorporated into the production line, access to technical data is often limited to online channels. However, researching online carries risks of leaking company information or other confidential data. Some companies use digital library management methods to manage online technical data, but these technologies are mostly based on text recognition and lack textual summarization, hindering subsequent display and retrieval.

[0057] The following explains the technical names used in this application:

[0058] A datasheet, also known as a "data manual" or "specification book," is an official document provided by the original manufacturer to the user, containing authoritative information such as the specifications, performance parameters, ordering information, and application examples of a component. A datasheet may include, for example, a brief introduction, brand, model, country of origin, schematic diagram, application scenarios, and a component reference table. The datasheet can be imported into a file server for storage using the import function.

[0059] Figure 1This is a schematic diagram illustrating the implementation environment of a product technical data management method according to an exemplary embodiment of this application. For example... Figure 1 As shown, the implementation environment includes terminal device 101, server 102 and server cluster 103.

[0060] in, Figure 1 The terminal device 101 shown may be, for example, a desktop computer, a tablet computer, a laptop computer, or any other terminal device that can be equipped with a text recognition model and an image recognition model, but is not limited thereto. Figure 1 The server 102 and server cluster 103 shown can be, for example, a server cluster or distributed system built on multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. There are no restrictions on these.

[0061] like Figure 1 As shown, the product technical data management method in this embodiment can be implemented, for example, through a terminal device 101 configured with a text recognition model and an image recognition model, by performing the following steps: acquiring product technical data; dividing the product technical data into regions based on the context outline in the product technical data to obtain multiple data regions, and extracting keywords to describe the content of the multiple data regions; dividing the multiple data regions into multiple text sub-regions and multiple image sub-regions based on the content of the multiple data regions, and converting the multiple text sub-regions into electronic text data and the multiple image sub-regions into image data; and constructing a relationship network based on the keywords, electronic text data, and image data.

[0062] It should be noted that, in this embodiment, product technical data can be obtained, for example, from server 102 and server cluster 103, or from a batch of multiple product technical data sets obtained from the network. This embodiment does not limit the method of obtaining product technical data. Furthermore, after obtaining electronic text data, image data, and relationship networks, these data can be stored in server 102 and server cluster 103. Server 102 and server cluster 103 can be, for example, internal company servers, and can be used for data storage and information retrieval.

[0063] To address the problems in existing technologies such as the lack of summarization of the content of each area in a book, the failure to associate the identification results with the initial data, and the disadvantages of display and retrieval, embodiments of this application propose a product technical data management method, a product technical data management device, an electronic device, a computer-readable storage medium, and a computer program product, which will be described in detail below.

[0064] Please see Figure 2 , Figure 2 This is a flowchart illustrating a product technical data management method in an exemplary embodiment of this application. This method can be applied to... Figure 1 The implementation environment is shown. It should be understood that this method can also be applied to other exemplary implementation environments and specifically executed by devices in other implementation environments. This embodiment does not limit the implementation environment to which the method is applicable.

[0065] like Figure 2 As shown, in an exemplary embodiment, the product technical data management method includes at least steps S210 to S240, which are described in detail below:

[0066] In step S210, product technical data is obtained.

[0067] For example, before obtaining product technical data, a pre-built client on terminal device 101 allows users to uniformly log in to the system for user authentication, group testing, or platform access. Whenever a new supplier or electronic component is available, the supplier provides the corresponding technical specifications and other information. The specifications include: introduction, brand, model, origin, schematic diagram, application scenarios, component comparison table, etc. These specifications are imported to server 102 and server cluster 103 for storage. Additionally, before R&D personnel conduct product development, they can download the specifications of the electronic components to be used and conventional electronic components from the network and store them on server 102 and server cluster 103.

[0068] It should be noted that, in the embodiments of this application, the product technical data includes not only electronic component specifications, but also design data, packaging libraries, and other technical data related to design, research and development, and production that can be retrieved on the Internet.

[0069] In step S220, the product technical data is divided into regions based on the context outline in the product technical data to obtain multiple data regions, and keywords used to describe the content of the multiple data regions are extracted.

[0070] In one embodiment of this application, a pre-built HED (Holistically-nested edge detection algorithm) is used to obtain the contextual outline of the product technical documentation, and the documentation is then divided into multiple data regions based on the contextual outline. When the product technical documentation is an electronic component specification, there are blank areas between different contents in the specification. These blank areas divide the specification into multiple content-focused regions, which may include, for example, an introductory content region, a test content region, and a table of contents region. The HED identifies the outlines between these content-focused regions, and then the regions are divided. In this embodiment, for example, a text-based sliding window-based prediction model language model NNLM (Neural Network Language Model) can be used, with the model interface trained by the algorithm platform, to dynamically extract keywords from each data region.

[0071] It should be noted that the subject terms in the embodiments of this application may include: catalog, pin definition, specification parameters, electrical characteristics, ordering information, package size, brand information, product category. Different subject terms describe the content of different data areas in the product technical data.

[0072] In step S230, the multiple data regions are divided according to their contents to obtain multiple text sub-regions and multiple image sub-regions. The multiple text sub-regions are then converted into electronic text data, and the multiple image sub-regions are converted into image data.

[0073] In this embodiment, a data region may contain multiple text sub-regions and image sub-regions. If text recognition or image recognition is directly performed on the data region, the text content will interfere with the image recognition result, or the image content will interfere with the text recognition result. In this embodiment, the multiple data regions are pre-divided according to the content of the multiple data regions to obtain multiple text sub-regions and image sub-regions, avoiding the problem of inaccurate recognition results caused by the interference of text content and image content with the recognition results in subsequent text recognition and other steps.

[0074] For each data area, when text content is present, a pre-built OCR (Optical Character Recognition) algorithm can be used to recognize the text, converting it into electronic text data and storing it in TIDB (a distributed NewSQL database). TIDB is built on server 102 and server cluster 103. For each data area, when image content is present, image data can be extracted using ROI (region of interest) and then stored in the image directory on server 102 and server cluster 103.

[0075] In step S240, a relationship network is constructed based on keywords, electronic text data, and image data.

[0076] For example, a directed relational network can be constructed using the extracted keywords, electronic text data, and image data. The content of this network may include, for example, the product technical document name or number, keywords, electronic text data links or descriptive statements, and image links or descriptive statements. Within the same product technical document, there is a mapping relationship between keywords, electronic text data, and image data, and also a mapping relationship between keywords and electronic text data and image data. By constructing this directed relational network, the relationship between the product technical document and the keywords, electronic text data, and image data can be visually displayed.

[0077] As can be seen from steps S210 to S240 above, the solution proposed in this embodiment divides regions and extracts keywords through the contextual outlines in the product technical data. The keywords can be used to summarize the content in multiple data regions. Dividing multiple data regions into multiple text sub-regions and image sub-regions can avoid mutual interference between text content and image content in the subsequent recognition process and improve the accuracy of recognition results. By constructing a relationship network through keywords, electronic text data and image data, the relationship structure can be displayed more intuitively, making it easier for relevant personnel to understand and retrieve.

[0078] In addition, the above methods can be used to download multiple specifications from the internet and other channels, digitize them, and store them in an internal database (such as an internal server). This avoids the leakage of company R&D goals and directions by leaving information when searching for keywords on the internet and other channels. Product technical data management is more convenient, and network security is also improved. Relevant technical personnel can directly search for relevant information in the internal database, avoiding the risk of accidentally entering phishing websites and leaking information when searching online.

[0079] In one embodiment of this application, Figure 2 Before dividing the product technical data into regions based on the context outline in the product technical data in step S220, the following steps are also included:

[0080] Input the electronic component specifications into the pre-built edge detection model;

[0081] The edge detection model is used to perform convolution and pooling on the multiple content concentration regions to obtain multiple initial contours of different sizes corresponding to different content concentration regions.

[0082] Based on the pre-trained weight parameters, multiple initial contours of different sizes corresponding to each content set region are weighted and fused to obtain context contours corresponding to multiple content set regions.

[0083] For example, the product technical data in this embodiment is an electronic component specification sheet. The electronic component specification sheet includes multiple content clusters separated by blank areas. These content clusters may contain, for example, text and / or image content. The electronic component specification sheet is input into a HED (Head-Up Display), and a neural network layer performs convolution and pooling on the multiple content clusters to obtain five initial contours of different sizes corresponding to each content cluster. These five initial contours are then weighted and fused to obtain the context contour. In this embodiment, context contour recognition avoids interference between different contents, ensuring the accuracy of subsequent image processing and text recognition results.

[0084] In one embodiment of this application, Figure 2 The step S220 shown above, which involves extracting subject terms to describe the content of the multiple data regions, includes the following steps:

[0085] If the content in the data area is text, then the data area is input into a pre-built keyword extraction model;

[0086] The topic word extraction model maps each word in the text to a feature vector, and merges the feature vectors of each word to obtain a multi-dimensional feature vector.

[0087] The probability of each word appearing in the multidimensional feature vector is calculated, and the topic words of the content of the multiple data regions are obtained according to the matching relationship between the preset topic words and the frequency of each word.

[0088] For example, when the content in a data area is text, keywords can be extracted to summarize the content. For instance, if the content is a brief introduction to the functions and dimensions of a product model, the keyword would be "introduction"; if the content describes functional testing of a product model in a specific experimental scenario, the keyword would be "testing". The data area is input into the NNLM (Neural Language Mapping Model), and the natural language is converted into feature vectors through a mapping matrix. In this embodiment, the feature vectors are word vectors. The word vectors of each word are merged to obtain a multi-dimensional feature vector. The probability of each word appearing in the multi-dimensional feature vector is calculated. By matching the probability with the keyword, the keyword for each data area can be identified.

[0089] In one embodiment of this application, in Figure 1 Step S230, as shown, converts the plurality of text sub-regions into electronic text data and the plurality of image sub-regions into image data, including the following steps:

[0090] The multiple text sub-regions are subjected to text recognition, the recognized text content is converted into an electronic text document, and a document number is assigned to the electronic text document to obtain the electronic text data;

[0091] The image content within the image sub-region is cropped according to the preset image size to obtain the image data.

[0092] For example, for each text sub-region, the text is recognized using an OCR text recognition algorithm and converted into an electronic text document. In this embodiment, the electronic text document is data content that can be read by computer devices such as computers. After obtaining the electronic text document, document numbers can be assigned to different electronic text documents. Different keywords correspond to different document numbers. During retrieval, once the keywords are obtained, the electronic text document can be directly extracted for selection and reading based on the correspondence between the keywords and document numbers. Compared to existing methods of searching through massive amounts of electronic text data using keywords, this application improves retrieval and display speed by assigning document numbers to electronic text documents.

[0093] It should be noted that when cropping the image content in a sub-region of an image according to the preset image size, the sizes of multiple cropped images are the same, and the image data is obtained by merging the multiple cropped images.

[0094] In one embodiment of this application, in Figure 1 Step S240, as shown, constructs a relationship network based on the keywords, electronic text data, and image data, including the following steps:

[0095] Obtain the first mapping relationship between keywords, electronic text data, and image data in the same product's technical documentation;

[0096] Based on the keywords, electronic text data, image data, and the first mapping relationship, construct a relationship network for the same product's technical information;

[0097] If multiple product technical documents contain the same product manufacturing address, then the relationship network of the multiple product technical documents is connected through the product manufacturing location.

[0098] For example, after performing the aforementioned region segmentation, text recognition, and image cropping operations on a product technical document, a mapping relationship exists between the keywords, electronic text data, and image data within the same product technical document. For instance, keyword 1 corresponds to electronic text data 1 and image data 1, and keyword 2 corresponds to electronic text data 2 and image data 2. In the relationship network, electronic text data and image data are represented using link addresses. Once the relationship network is constructed, it can be stored in the Neo4j database (a high-performance web-oriented database). When two product technical documents share the same production address, their relationship networks can be connected via the production address, facilitating the determination of whether the two documents belong to the same file family. When two product technical documents belong to the same file family, the relationship degree is obtained through the distance between elements; the relationship degree is the distance between the two product technical documents.

[0099] See Figure 3 , Figure 3 This is a schematic diagram of a relational network shown in an exemplary embodiment of this application. Figure 3 In the data structure, the relationship network of Specification 1 includes Image Link 1, Introduction 1, Description 1, and Test 1, while the relationship network of Specification 2 includes Image Link 2, Introduction 2, Description 2, and Test 2. The product manufacturing addresses in Specification 1 and Specification 2 are the same, namely, Manufacturing Location 1. In this case, the relationship degree between Specification 1 and Specification 2 is the distance between them, which is 2.

[0100] In one embodiment of this application, in Figure 1 After constructing the relationship network based on the keywords, electronic text data, and image data in step S240, the following steps are also included:

[0101] Obtain the product model from the product technical data;

[0102] All characters in the electronic text data are segmented to obtain multiple words;

[0103] A second mapping relationship is constructed between the multiple terms, product models, subject terms, and document numbers in the same product technical documents, and the multiple terms, product models, subject terms, and document numbers are stored to obtain a retrieval database.

[0104] For example, product model information can be obtained from product technical documents using keyword retrieval technology. The jieba word segmentation toolkit (a Chinese word segmentation library) is used to segment all text in all sentences or paragraphs of the electronic text data, resulting in multiple terms. Multiple terms from multiple product technical documents correspond to different product models, and multiple terms from the same product technical document correspond to different subject terms and document numbers. Therefore, a second mapping relationship needs to be constructed between these multiple terms, product models, subject terms, and document numbers within the same product technical document. Then, these multiple terms, product models, subject terms, and document numbers are stored in the TIDB database to obtain the retrieval database.

[0105] In one embodiment of this application, after obtaining the search database, the following steps are further included:

[0106] Obtain search keywords, which include search topic terms, and / or search product models, and / or search terms. The search priority of search product models is higher than that of search topic terms, and the search priority of search topic terms is higher than that of search terms.

[0107] Based on the search keywords and their priority, information in the search database is retrieved to obtain initial search results, which include product model, and / or subject terms, and / or multiple words, and / or document number.

[0108] Obtain the number of times the product model appears in the initial search results, and / or the number of times the keyword appears, and / or the number of times multiple words appear, and / or the number of times the document number appears;

[0109] The display order of the contents in the initial search results is sorted according to the frequency of occurrence, and the electronic text data corresponding to the initial search results is obtained according to the second mapping relationship.

[0110] For example, the system obtains the search information entered by the user on the client side. This search information may include multiple search keywords, each with a different priority; higher-priority keywords are retrieved first. The system then searches the database based on these keywords to obtain initial search results. These initial results contain multiple product models, and / or keywords, and / or multiple terms, and / or document numbers corresponding to the keywords. The order of displaying the search content is determined by analyzing the frequency of occurrence of product models, and / or keywords, and / or multiple terms, and / or document numbers. For example, if the keyword "test" appears most frequently, then electronic text data related to "test" should be displayed first. When displaying data from the database, the corresponding electronic text data can be determined based on the second mapping relationship constructed above. This data is then passed to a PDF (Portable Document Format) object display box on the Vue (a progressive front-end framework)-based front-end page via a Java (object-oriented programming language) backend program.

[0111] It should be noted that when the keywords, terms, and document numbers in the search results belong to different product technical documents, the distance between multiple product technical documents can be determined through the relationship network stored in neo4j. When the distance does not exceed the preset distance threshold, the electronic text data of multiple product technical documents can be displayed.

[0112] In one embodiment of this application, the above method can be applied to a specification digitization management system to digitize electronic specifications and construct a retrieval database. See also Figure 4 , Figure 4 This is an exemplary embodiment of the specification digital management system and its interface architecture with external systems. The specification digital management system provides query and retrieval capabilities to systems such as algorithm platform, R&D platform, procurement platform, and production platform through an external service interface module via an intranet gateway and unified authentication, using fixed commands. It can seamlessly connect to other applications.

[0113] Figure 5 This is a block diagram illustrating a product technical data management device according to an exemplary embodiment of this application. The device can be applied to… Figure 1 The implementation environment shown is not limited to this embodiment. This device can also be applied to other exemplary implementation environments and specifically configured in other devices. This embodiment does not limit the implementation environment to which the device is applicable.

[0114] like Figure 5 As shown, the exemplary product technical data management device includes:

[0115] Data acquisition module 501 is used to acquire product technical data;

[0116] The keyword acquisition module 502 is used to divide the product technical data into regions based on the context outline in the product technical data to obtain multiple data regions, and extract keywords to describe the content of the multiple data regions.

[0117] The electronic text data and image data acquisition module 503 is used to divide the multiple data regions according to the content of the multiple data regions to obtain multiple text sub-regions and multiple image sub-regions, and convert the multiple text sub-regions into electronic text data and the multiple image sub-regions into image data;

[0118] The relationship network construction module 504 is used to construct a relationship network based on the keywords, electronic text data and image data. The relationship network is used to display the relationship between the product technical information and the keywords, electronic text data and image data.

[0119] In this exemplary product technical data management device, product technical data is divided into regions, and keywords are extracted from each region. A relationship network is constructed based on the relationship between these keywords and the content of the product technical data. This not only digitizes the product technical data but also improves its display and retrieval capabilities.

[0120] In another exemplary embodiment, the product technical data management device further includes a context profile acquisition module, which is mainly used for:

[0121] Input the electronic component specifications into the pre-built edge detection model;

[0122] The edge detection model is used to perform convolution and pooling on the multiple content concentration regions to obtain multiple initial contours of different sizes corresponding to different content concentration regions.

[0123] Based on the pre-trained weight parameters, multiple initial contours of different sizes corresponding to each content set region are weighted and fused to obtain context contours corresponding to multiple content set regions.

[0124] In another exemplary embodiment, the keyword acquisition module 502 includes a keyword acquisition unit, which is configured to:

[0125] If the content in the data area is text, then the data area is input into a pre-built keyword extraction model;

[0126] The topic word extraction model maps each word in the text to a feature vector, and merges the feature vectors of each word to obtain a multi-dimensional feature vector.

[0127] The probability of each word appearing in the multidimensional feature vector is calculated, and the topic words of the content of the multiple data regions are obtained according to the matching relationship between the preset topic words and the frequency of each word.

[0128] In another exemplary embodiment, the electronic text data and image data acquisition module 503 includes a sub-region content processing unit, which is configured to:

[0129] The multiple text sub-regions are subjected to text recognition, the recognized text content is converted into an electronic text document, and a document number is assigned to the electronic text document to obtain the electronic text data;

[0130] The image content within the image sub-region is cropped according to the preset image size to obtain the image data.

[0131] In another exemplary embodiment, the relationship network building module 504 is configured to:

[0132] Obtain the first mapping relationship between keywords, electronic text data, and image data in the same product's technical documentation;

[0133] Based on the keywords, electronic text data, image data, and the first mapping relationship, construct a relationship network for the same product's technical information;

[0134] If multiple product technical documents contain the same product manufacturing address, then the relationship network of the multiple product technical documents is connected through the product manufacturing location.

[0135] In another exemplary embodiment, the product technical data management device further includes a search database construction module, which is configured to:

[0136] Obtain the product model from the product technical data;

[0137] All characters in the electronic text data are segmented to obtain multiple words;

[0138] A second mapping relationship is constructed between the multiple terms, product models, subject terms, and document numbers in the same product technical documents, and the multiple terms, product models, subject terms, and document numbers are stored to obtain a retrieval database.

[0139] In another exemplary embodiment, the product technical data management device further includes a retrieval module, which is configured to:

[0140] Obtain search keywords, which include search topic terms, and / or search product models, and / or search terms. The search priority of search product models is higher than that of search topic terms, and the search priority of search topic terms is higher than that of search terms.

[0141] Based on the search keywords and their priority, information in the search database is retrieved to obtain initial search results, which include product model, and / or subject terms, and / or multiple words, and / or document number.

[0142] Obtain the number of times the product model appears in the initial search results, and / or the number of times the keyword appears, and / or the number of times multiple words appear, and / or the number of times the document number appears;

[0143] The display order of the contents in the initial search results is sorted according to the frequency of occurrence, and the electronic text data corresponding to the initial search results is obtained according to the second mapping relationship.

[0144] It should be noted that the product technical data management device and the product technical data management method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments and will not be repeated here. In practical applications, the product technical data management device provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0145] Embodiments of this application also provide an electronic device, including: one or more processors; and a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the product technical data management method provided in the above embodiments.

[0146] Figure 6 A schematic diagram of a computer system suitable for an electronic device according to an embodiment of this application is shown. It should be noted that... Figure 6 The computer system 600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0147] like Figure 6As shown, the computer system 600 includes a Central Processing Unit (CPU) 601, which can perform various appropriate actions and processes based on programs stored in Read-Only Memory (ROM) 602 or programs loaded from Storage Unit 608 into Random Access Memory (RAM) 603, such as performing the methods described in the above embodiments. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An Input / Output (I / O) interface 605 is also connected to the bus 604.

[0148] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.

[0149] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs various functions defined in the system of this application.

[0150] It should be noted that the computer-readable medium shown in the embodiments of this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wired, etc., or any suitable combination thereof.

[0151] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. Each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0152] The units described in the embodiments of this application can be implemented in software or hardware, and the described units can also be located in a processor. The names of these units do not necessarily limit the specific unit itself.

[0153] Another aspect of this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the product technical data management method as described above. This computer-readable storage medium may be included in the electronic device described in the above embodiments, or it may exist independently and not assembled into the electronic device.

[0154] Another aspect of this application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the product technical data management method provided in the various embodiments described above.

[0155] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A method for managing product technical data, characterized in that, The method includes: Obtain product technical information; Based on the contextual outline in the product technical data, the product technical data is divided into regions to obtain multiple data regions, and keywords used to describe the content of the multiple data regions are extracted. The multiple data regions are divided according to their contents to obtain multiple text sub-regions and multiple image sub-regions. The multiple text sub-regions are then converted into electronic text data, and the multiple image sub-regions are converted into image data. A relationship network is constructed based on the keywords, electronic text data, and image data. This relationship network is used to demonstrate the relationship between the product technical information and the keywords, electronic text data, and image data.

2. The product technical data management method according to claim 1, characterized in that, The product technical documentation includes electronic component specifications, which comprise multiple spaced-apart content areas. Before dividing the product technical documentation into areas based on the contextual outlines within the documentation, the process further includes: Input the electronic component specifications into the pre-built edge detection model; The edge detection model is used to perform convolution and pooling on the multiple content concentration regions to obtain multiple initial contours of different sizes corresponding to different content concentration regions. Based on the pre-trained weight parameters, multiple initial contours of different sizes corresponding to each content set region are weighted and fused to obtain the context contours corresponding to multiple content set regions.

3. The product technical data management method according to claim 1, characterized in that, The extraction of subject terms used to describe the content of the multiple data regions includes: If the content in the data area is text, then the data area is input into a pre-built keyword extraction model; The topic word extraction model maps each word in the text to a feature vector, and merges the feature vectors of each word to obtain a multi-dimensional feature vector. The probability of each word appearing in the multidimensional feature vector is calculated, and the topic words of the content of the multiple data regions are obtained according to the matching relationship between the preset topic words and the frequency of each word.

4. The product technical data management method according to claim 1, characterized in that, The process of converting the plurality of text sub-regions into electronic text data and the plurality of image sub-regions into image data includes: The multiple text sub-regions are subjected to text recognition, the recognized text content is converted into an electronic text document, and a document number is assigned to the electronic text document to obtain the electronic text data; The image content within the image sub-region is cropped according to the preset image size to obtain the image data.

5. The product technical data management method according to claim 1, characterized in that, The step of constructing a relationship network based on the keywords, electronic text data, and image data includes: Obtain the first mapping relationship between keywords, electronic text data, and image data in the same product's technical documentation; Based on the keywords, electronic text data, image data, and the first mapping relationship, construct a relationship network for the same product's technical information; If multiple product technical documents contain the same product manufacturing address, then the relationship network of the multiple product technical documents is connected through the product manufacturing location.

6. The product technical data management method according to claim 4, characterized in that, After constructing the relationship network based on the aforementioned keywords, electronic text data, and image data, the process further includes: Obtain the product model from the product technical data; All characters in the electronic text data are segmented to obtain multiple words; A second mapping relationship is constructed between the multiple terms, product models, subject terms, and document numbers in the same product technical documents, and the multiple terms, product models, subject terms, and document numbers are stored to obtain a retrieval database.

7. The product technical data management method according to claim 6, characterized in that, After obtaining the search database, the process also includes: Obtain search keywords, which include search topic terms, and / or search product models, and / or search terms. The search priority of search product models is higher than that of search topic terms, and the search priority of search topic terms is higher than that of search terms. Based on the search keywords and their priority, information in the search database is retrieved to obtain initial search results, which include product model, and / or subject terms, and / or multiple words, and / or document number. Obtain the number of times the product model appears in the initial search results, and / or the number of times the keyword appears, and / or the number of times multiple words appear, and / or the number of times the document number appears; The display order of the contents in the initial search results is sorted according to the frequency of occurrence, and the electronic text data corresponding to the initial search results is obtained according to the second mapping relationship.

8. A product technical data management device, characterized in that, The device includes: The data acquisition module is used to acquire product technical data; The keyword acquisition module is used to divide the product technical data into regions based on the context outline in the product technical data, obtain multiple data regions, and extract keyword terms to describe the content of the multiple data regions. The electronic text data and image data acquisition module is used to divide the multiple data regions according to the content of the multiple data regions to obtain multiple text sub-regions and multiple image sub-regions, and convert the multiple text sub-regions into electronic text data and the multiple image sub-regions into image data; The relationship network construction module is used to construct a relationship network based on the keywords, electronic text data, and image data. The relationship network is used to display the relationship between the product technical information and the keywords, electronic text data, and image data.

9. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to implement the product technical data management method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the product technical data management method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Digital book structured analysis processing method

    CN106250830A

  • Book storage method, system and device based on OCR (Optical Character Recognition) technology

    CN112905733A

  • Method for automatically classifying electronic archives based on deep learning

    CN109658062A

  • Document processor and document processing method

    JP2003288334A