Data processing method, apparatus, device, and medium
By using predefined industry rules for keyword matching and entity labeling in data processing, combined with correlation analysis of user behavior data, the problem of long model building cycles and low efficiency in existing technologies is solved, and efficient business data analysis is achieved.
Patent Information
- Application Number
- CN202211493507.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-25
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2042-11-25
AI Technical Summary
Existing technologies require the construction of entity recognition models specifically for the target industry when conducting industry trend analysis, consumer intent analysis, and brand competition analysis. This results in long model construction cycles, low efficiency in business data analysis, and high labor costs.
By using predefined industry rules, keyword matching and entity labeling are performed on the data to be analyzed, and statistical analysis is conducted based on the correlation of user behavior data, thus eliminating the need for training entity recognition models.
It improved the efficiency of industry business data analysis, saved labor costs, and increased the speed and accuracy of analysis.
Smart Images

Figure CN115759100B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to the field of data processing and big data, specifically to a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] Industry / brand / product analysis and insights are crucial throughout the pre-launch, during-launch, and post-launch phases of marketing campaigns. Before launch, industry and brand trends can be analyzed to explore potential advertising opportunities. During launch, industry / brand insights can be used to select specific target audiences for targeted advertising. Post-launch, further analysis of brand trends allows for the measurement and attribution of post-launch effectiveness.
[0003] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0004] This disclosure provides a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product.
[0005] According to one aspect of this disclosure, a data processing method is provided, comprising: acquiring at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; acquiring at least one preset rule, wherein each preset rule includes at least one entity keyword; performing keyword matching on each piece of data to be analyzed based on each entity keyword in the at least one preset rule to determine at least one entity tag of the data to be analyzed, wherein each entity tag corresponds to the matched entity keyword; aggregating user behavior data associated with each piece of data to be analyzed based on the at least one entity tag of each piece of data to be analyzed to obtain a set of user behavior data corresponding to each entity tag; and performing statistical analysis based on the set of user behavior data corresponding to each entity tag to obtain business data analysis results.
[0006] According to another aspect of this disclosure, a data processing apparatus is provided, comprising: a first acquisition unit configured to acquire at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; a second acquisition unit configured to acquire at least one preset rule, wherein each preset rule includes at least one entity keyword; a matching unit configured to perform keyword matching on each piece of data to be analyzed based on each entity keyword in the at least one preset rule to determine at least one entity tag of the data to be analyzed, wherein each entity tag corresponds to the matched entity keyword; an aggregation unit configured to aggregate user behavior data associated with each piece of data to be analyzed based on at least one entity tag of each piece of data to be analyzed to obtain a set of user behavior data corresponding to each entity tag; and an analysis unit configured to perform statistical analysis based on the set of user behavior data corresponding to each entity tag to obtain business data analysis results.
[0007] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the data processing method described above.
[0008] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the above-described data processing method.
[0009] According to another aspect of this disclosure, a computer program product is provided, including a computer program, wherein the computer program implements the above-described data processing method when executed by a processor.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0012] Figure 1A schematic diagram of an exemplary system in which the various methods described herein may be implemented according to embodiments of the present disclosure is shown;
[0013] Figure 2 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;
[0014] Figure 3 A flowchart of a data processing method according to an embodiment of the present disclosure is shown;
[0015] Figure 4 A structural block diagram of a data processing apparatus according to an embodiment of the present disclosure is shown;
[0016] Figure 5 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0018] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0019] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0020] In related technologies, when conducting business data analysis on a target industry, such as industry trend analysis, consumer intent analysis, brand competition analysis, and consumer decision analysis, it is usually necessary to first build an entity recognition model specifically for that target industry to identify entities such as industry categories and market segments. During model building, a large amount of manual data annotation is often required, and the model also needs to undergo multiple iterations, resulting in a long model building cycle and low efficiency in business data analysis.
[0021] According to embodiments of this disclosure, keyword matching can be performed on the data to be analyzed based on predefined industry rules, thereby assigning at least one entity label to each data point. Based on the correlation between the data to be analyzed and user behavior data, statistical analysis is performed according to different entity label dimensions to obtain the business data analysis results for the industry. Thus, the analysis of industry data can be completed without training an entity recognition model, thereby improving the efficiency of industry business data analysis and saving labor costs.
[0022] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0023] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0024] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of the data processing methods described above.
[0025] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105, and / or 106 under a Software as a Service (SaaS) model.
[0026] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 This is an example of a system used to implement the various methods described herein, and is not intended to be limiting.
[0027] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to acquire data to be analyzed. The client devices can provide an interface that allows users to interact with them. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0028] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0029] Network 110 can be any type of network well known to those skilled in the art, and can use any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0030] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0031] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0032] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0033] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0034] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located away from server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0035] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0036] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0037] According to some embodiments, such as Figure 2 As shown, a data processing method is provided, including: step S201, acquiring at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; step S202, acquiring at least one preset rule, wherein each preset rule includes at least one entity keyword; step S203, performing keyword matching on each piece of data to be analyzed based on each entity keyword in the at least one preset rule to determine at least one entity tag of the data to be analyzed, wherein each entity tag corresponds to the matched entity keyword; step S204, aggregating user behavior data associated with each piece of data to be analyzed based on at least one entity tag of each piece of data to be analyzed to obtain a user behavior data set corresponding to each entity tag; and step S205, performing statistical analysis on the user behavior data set corresponding to each entity tag to obtain business data analysis results.
[0038] According to embodiments of this disclosure, keyword matching can be performed on the data to be analyzed based on predefined industry rules, thereby assigning at least one entity label to each data point. Based on the correlation between the data to be analyzed and user behavior data, statistical analysis is performed according to different entity label dimensions to obtain the business data analysis results for the industry. Thus, the analysis of industry data can be completed without training an entity recognition model, thereby improving the efficiency of industry business data analysis and saving labor costs.
[0039] In some embodiments, the data to be analyzed may be at least one of user search text and webpage title.
[0040] In some embodiments, the user search text may be the search request text entered by the user into the search engine. The webpage title may include the title information of the resource (e.g., webpage, video, image, etc.) pointed to by any URL (Uniform Resource Locator) address on the Internet.
[0041] In some embodiments, keyword matching and data annotation can be performed based on the full amount of search request text and webpage titles that can be obtained.
[0042] In some embodiments, the scope of the data to be analyzed can be defined before labeling the data to be analyzed.
[0043] In some embodiments, obtaining user search text from at least one set of data to be analyzed may include: obtaining a plurality of first user search texts; performing industry category prediction on each of the plurality of first user search texts to determine the industry category to which each first user search text belongs; and obtaining at least one user search text from the plurality of first user search texts whose industry category is the target industry category, as at least one user search text in at least one set of data to be analyzed.
[0044] Therefore, before entity annotation, the user search text in the data to be analyzed is filtered and the scope is defined, thereby filtering the data in the initial stage, avoiding unnecessary calculations, saving computing resources, and improving analysis efficiency.
[0045] In some embodiments, industry category prediction can be performed on each user search text in the full set of user search texts to filter out the user search texts that correspond to the target industry as data to be analyzed.
[0046] In some embodiments, industry category prediction can be performed on each user search text in the full set of user search texts within a preset time range, thereby filtering out the user search texts corresponding to the target industry as data to be analyzed. This ensures the timeliness of the data to be analyzed.
[0047] In some embodiments, the daily new user search texts can be determined through comparison, and then only the industry category prediction of these new user search texts can be performed to filter out the user search texts corresponding to the target industry as data to be analyzed. This avoids a large amount of redundant calculations, saves computing resources, and improves computational efficiency.
[0048] In some embodiments, the industry category prediction described above can be implemented, for example, based on a pre-trained industry classification model. In some embodiments, this model can be built on, for example, a convolutional neural network and can be trained using sample text labeled with industry category tags.
[0049] In some embodiments, the format of user search text can be specified to filter user search text. For example, the format of user search text can be defined as "user search text containing 'vehicle'", "user search text containing 'brand A'", etc.
[0050] In some embodiments, obtaining the webpage title from at least one set of data to be analyzed may include: obtaining at least one target site, wherein the content category of each target site corresponds to a target industry category; and extracting the webpage title contained in each webpage of the at least one target site as at least one webpage title in at least one set of data to be analyzed.
[0051] Therefore, before entity annotation, the webpage titles in the data to be analyzed are filtered and the scope is defined, thereby filtering the data in the initial stage, avoiding unnecessary calculations, saving computing resources, and improving analysis efficiency.
[0052] In some embodiments, one or more target websites under the target industry category can be obtained first, and the title information (i.e., webpage title) of the resource pointed to by each URL address in the target websites can be extracted, thereby defining the scope of the data to be analyzed, such as webpage titles.
[0053] In some embodiments, after determining at least one piece of data to be analyzed, at least one preset rule may be further obtained.
[0054] In some embodiments, at least one preset rule may be a preset rule targeting the same target industry, and the preset rules may constitute a rule set corresponding to that target industry. Each preset rule includes a topic name and at least one entity keyword corresponding to that topic name.
[0055] In some embodiments, at least one entity keyword can be further categorized into phrase matching keywords, exact matching keywords, etc. In some embodiments, different keyword matching methods can be applied to different entity keyword types.
[0056] For example, a preset rule can be: {"humidifier":{"phrase matching keyword":["humidifier"],"exact match keyword":[],"negative keyword":[]},"dehumidifier":{"phrase matching keyword":["dehumidifier"],"exact match keyword":[],"negative keyword":[]},"key":"market segment","air purifier":{"phrase matching keyword":["air purifier"],"exact match keyword":[],"negative keyword":[]}}.
[0057] The theme of this preset rule is "market segmentation". The entity tags of the multiple market segments in this preset rule include "humidifier", "dehumidifier" and "purifier", and corresponding phrase matching keywords and exact matching keywords are defined for each entity tag.
[0058] In some embodiments, keyword matching for phrase matching can be performed through semantic matching. Semantic matching refers to inputting the phrase matching keyword and the data to be analyzed into a pre-trained language model, obtaining their respective semantic codes, and considering the data to be analyzed as matching the phrase matching keyword if the similarity between the two semantic codes is less than a preset similarity threshold. Then, the entity label corresponding to the phrase matching keyword can be labeled on the data to be analyzed.
[0059] In some embodiments, exact match keywords can be matched using literal matching, for example. Literal matching can be based on a Bag-of-words model to encode the exact match keyword and multiple word segments in the data to be analyzed, and then match them. When there is a word segment in the data to be analyzed that exactly matches the keyword, the entity tag corresponding to the exact match keyword can be labeled on the data to be analyzed.
[0060] In some embodiments, multiple preset rules may correspond to different topic names, and there may be hierarchical relationships between multiple topic names.
[0061] In some exemplary embodiments, multiple topic names may include industry, market segment, brand, product, focus, etc. Here, brand is the parent topic of product, market segment is the parent topic of brand, and industry is the parent topic of market segment. Correspondingly, the entity tags in entity rules at different levels also have corresponding hierarchical relationships. For example, the entity tag "Brand A" includes multiple subordinate entity tags such as "Product a", "Product b", and "Product c".
[0062] In some embodiments, the number of at least one preset rules can be multiple, and the first rule and the second rule in the at least one preset rule have an entity hierarchical relationship, and the entity type of the first rule is the superior type of the entity type of the second rule. Determining at least one entity tag of the data to be analyzed further includes: in response to the fact that the at least one entity tag of the data to be analyzed includes an entity tag corresponding to the entity keyword in the second rule, based on the entity hierarchical relationship between the first rule and the second rule, determining the entity tag of the data to be analyzed that corresponds to the corresponding entity keyword in the first rule.
[0063] In some exemplary embodiments, when keyword matching is performed on data to be analyzed and only the entity tag "Product A" is matched, the parent entity tag "Brand A", the parent entity tag "Sub-product 1" of "Brand A", and the parent entity tag "Industry 1" of "Sub-product 1" can be annotated on the data to be analyzed based on the aforementioned hierarchical relationship between entities. Thus, based on the hierarchical relationship between keywords, more comprehensive entity annotation is achieved on the data to be analyzed, thereby improving the accuracy and comprehensiveness of subsequent business data analysis and statistics.
[0064] In some embodiments, the first preset rule in at least one preset rule includes a negative keyword. Performing keyword matching on each piece of data to be analyzed in at least one piece of data to determine at least one entity label of the data to be analyzed includes: performing keyword matching on the data to be analyzed based on each entity keyword in the first preset rule to determine at least one candidate keyword of the data to be analyzed; and in response to the inclusion of a negative keyword in at least one candidate keyword, filtering out the corresponding keyword in at least one candidate keyword to determine at least one entity label of the data to be analyzed based on the remaining keywords in at least one candidate keyword.
[0065] Therefore, by filtering negative keywords during the annotation process, the accuracy of entity annotation can be further improved.
[0066] In some embodiments, one or more preset rules may include negative keywords. During the keyword matching process, after the corresponding keyword is matched, the matched candidate keywords are filtered based on the negative keyword matching, thereby avoiding some mislabeling and further improving the accuracy of entity labeling.
[0067] In some embodiments, such as Figure 3 As shown, the data processing method may further include: for each piece of data to be analyzed in at least one set of data to be analyzed: step S301, segmenting the data to be analyzed into multiple words; step S302, performing named entity recognition on the multiple words to obtain the entity type of each word in the multiple words; and step S303, in response to at least two words in the multiple words having the same entity type, labeling the at least two words as co-entities; step S304, in response to the statistical probability of the first co-entity in at least one set of data to be analyzed being greater than a preset probability threshold, the first word in the first co-entity does not match the entity keyword in at least one preset rule, and the second word in the first co-entity matches the entity keyword in at least one preset rule, adding the first word as an entity keyword to the corresponding preset rule of the second word to update at least one preset rule; and step S305, performing keyword matching on each piece of data to be analyzed in at least one set of data to be analyzed based on the updated at least one preset rule to update the entity label of the data to be analyzed.
[0068] Therefore, based on the highly co-occurring entities in the data to be analyzed, the keywords not covered in the initial rule set are expanded, so that the business data of the target industry can be more comprehensively mined and analyzed without training the entity recognition model of the industry.
[0069] In some embodiments, during the entity annotation process of the data to be analyzed, co-existing entities can also be statistically analyzed. Co-existing entities can be multiple keywords with the same named entity type appearing in the same data to be analyzed. For example, if a user's search text is "humidifier comparison between brand A and brand B", and after segmenting the user's search text into words, and then performing named entity recognition on multiple word segments, if the named entity type of both "brand A" and "brand B" is "brand", then "brand A" and "brand B" are a pair of co-existing entities.
[0070] In some embodiments, probability statistics can be performed on co-existing entities similar to those described above in all the data to be analyzed. If the statistical probability of a co-existing entity is greater than a preset probability threshold, and one or more of the co-existing entities are included in the rule set, but the co-existing entity also contains keywords not belonging to the rule set, then the keyword and its corresponding entity tag can be added to the corresponding rule. In some embodiments, after adding the entity tag, rules for the lower-level entities of that entity tag can be further defined.
[0071] The aforementioned named entity recognition can be performed based on a pre-trained named entity recognition model. This named entity recognition model differs from the industry-specific entity recognition models mentioned in the related technologies above; instead, it is a general-purpose named entity model applicable to various industries, thus eliminating the need for separate training based on different target industries.
[0072] Subsequently, the data to be analyzed can be re-labeled based on the updated rule set, thereby obtaining more complete labeling information for each piece of data to be analyzed.
[0073] Each piece of data to be analyzed is associated with corresponding user behavior data. For example, for a user's search text, the relevant user behavior data may include the user ID, the time, location, device ID, and context of the user's search; for each webpage title, the relevant user behavior data may include the user ID who viewed the page and the time the user stayed on the page.
[0074] Based on the labeled data to be analyzed, user behavior data can be aggregated according to each entity label to obtain a set of user behavior data associated with each entity label. Then, based on the user behavior data set and the corresponding entity labels, statistical analysis of user behavior data can be performed using appropriate data analysis and statistical methods to obtain corresponding business data analysis results.
[0075] Understandably, the data analysis and statistical methods described above can be determined based on actual needs. For example, they may include industry trend analysis, consumer intent analysis, brand competition analysis, consumer decision analysis, etc., without any restrictions.
[0076] In some embodiments, such as Figure 4As shown, a data processing apparatus 400 is provided, comprising: a first acquisition unit 410 configured to acquire at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; a second acquisition unit 420 configured to acquire at least one preset rule, wherein each preset rule includes at least one entity keyword; a matching unit 430 configured to perform keyword matching on each piece of data to be analyzed based on each entity keyword in the at least one preset rule to determine at least one entity tag of the data to be analyzed, wherein each entity tag corresponds to the matched entity keyword; an aggregation unit 440 configured to aggregate user behavior data associated with each piece of data to be analyzed based on at least one entity tag of each piece of data to be analyzed to obtain a set of user behavior data corresponding to each entity tag; and an analysis unit 450 configured to perform statistical analysis based on the set of user behavior data corresponding to each entity tag to obtain business data analysis results.
[0077] The operations performed by units 410-450 in the data processing device 400 are similar to the operations of steps S201 to S205 in the above data processing method, and will not be described in detail here.
[0078] In some embodiments, the data processing apparatus may further include: an execution unit configured to perform the operations of the following sub-units for each of the at least one set of data to be analyzed, the execution unit including: a word segmentation sub-unit configured to segment the data to be analyzed into multiple words; an identification sub-unit configured to perform named entity recognition on the multiple words to obtain the entity type of each of the multiple words; and a labeling sub-unit configured to label at least two words as co-entity entities in response to at least two words having the same entity type; an update unit configured to add the first word as an entity keyword to the corresponding preset rule of the second word in response to a statistical probability greater than a preset probability threshold in at least one set of data, a first word in the first co-entity entity not matching an entity keyword in at least one preset rule, and a second word in the first co-entity entity matching an entity keyword in at least one preset rule, thereby updating at least one preset rule; and a matching unit further configured to: perform keyword matching on each of the at least one set of data to be analyzed based on the updated at least one preset rule, thereby updating the entity label of the data to be analyzed.
[0079] In some embodiments, the number of at least one preset rules is multiple, the first rule and the second rule in the at least one preset rules have an entity hierarchy relationship, and the entity type of the first rule is the parent type of the entity type of the second rule. The matching unit further includes: a determining subunit, configured to determine the entity tag of the data to be analyzed that corresponds to the entity keyword in the first rule based on the entity hierarchy relationship between the first rule and the second rule in response to at least one entity tag in the data to be analyzed including an entity tag corresponding to the entity keyword in the second rule.
[0080] In some embodiments, the first preset rule in at least one preset rule includes a negative keyword, and the matching unit includes: a matching subunit configured to perform keyword matching on the data to be analyzed based on each entity keyword in the first preset rule to determine at least one candidate keyword in the data to be analyzed; and a filtering subunit configured to filter out the corresponding keyword in at least one candidate keyword in response to the inclusion of a negative keyword in at least one candidate keyword, so as to determine at least one entity label of the data to be analyzed based on the remaining keywords in at least one candidate keyword.
[0081] In some embodiments, the first acquisition unit includes: a first acquisition subunit configured to acquire a plurality of first user search texts; a prediction subunit configured to predict the industry category of each of the plurality of first user search texts to determine the industry category to which each first user search text belongs; and a second acquisition subunit configured to acquire at least one user search text among the plurality of first user search texts whose industry category is the target industry category, as at least one user search text in at least one set of data to be analyzed.
[0082] In some embodiments, the first acquisition unit includes: a third acquisition subunit configured to acquire at least one target site, wherein the content category of each target site corresponds to a target industry category; and an extraction subunit configured to extract the webpage title contained in each webpage of the at least one target site as at least one webpage title in at least one set of data to be analyzed.
[0083] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0084] According to embodiments of this disclosure, an electronic device, a readable storage medium, and a computer program product are also provided.
[0085] refer to Figure 5The present invention describes a structural block diagram of an electronic device 500 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0086] like Figure 5 As shown, the electronic device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0087] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, output unit 507, storage unit 508, and communication unit 509. Input unit 506 can be any type of device capable of inputting information to electronic device 500. Input unit 506 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 507 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 508 may include, but is not limited to, disk and optical disk. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0088] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as the data processing methods described above. For example, in some embodiments, the data processing methods described above can be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the data processing methods described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the data processing methods described above by any other suitable means (e.g., by means of firmware).
[0089] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0090] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0091] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0092] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0093] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0094] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0095] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0096] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A data processing method, the method comprising: Obtain at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; Obtain at least one preset rule, wherein each preset rule includes at least one entity keyword; Based on each entity keyword in the at least one preset rule, keyword matching is performed on each piece of data to be analyzed in the at least one piece of data to be analyzed, so as to determine at least one entity tag of the data to be analyzed, wherein each entity tag in the at least one entity tag corresponds to the matched entity keyword; Based on at least one entity label of each of the at least one data to be analyzed, the user behavior data associated with each of the at least one data to be analyzed is aggregated to obtain a set of user behavior data corresponding to each entity label; and Statistical analysis is performed on the user behavior data set corresponding to each entity tag to obtain business data analysis results.
2. The method according to claim 1, further comprising: For each of the at least one data points to be analyzed: The data to be analyzed is segmented into words to obtain multiple word segments; Named entity recognition is performed on the multiple word segments to obtain the entity type of each word segment. as well as In response to the fact that at least two of the multiple word segments have the same entity type, the at least two word segments are labeled as common entity types; In response to the statistical probability of a first co-occurring entity in the at least one piece of data to be analyzed being greater than a preset probability threshold, a first word segment in the first co-occurring entity not matching an entity keyword in the at least one preset rule, and a second word segment in the first co-occurring entity matching an entity keyword in the at least one preset rule, the first word segment is added as an entity keyword to the preset rule corresponding to the second word segment to update the at least one preset rule; and Based on at least one updated preset rule, keyword matching is performed on each piece of data to be analyzed in the at least one piece of data to be analyzed, so as to update the entity label of the data to be analyzed.
3. The method according to claim 1 or 2, wherein, The number of the at least one preset rule is multiple, and the first rule and the second rule in the at least one preset rule have a hierarchical relationship in terms of entities, and the entity type of the first rule is the parent type of the entity type of the second rule. Determining at least one entity label of the data to be analyzed further includes: In response to the fact that at least one entity tag in the data to be analyzed includes an entity tag corresponding to an entity keyword in the second rule, the entity tag corresponding to the entity keyword in the first rule in the data to be analyzed is determined based on the hierarchical relationship between the entities in the first rule and the second rule.
4. The method according to claim 1 or 2, wherein, The first preset rule in the at least one preset rule includes a negative keyword, and the step of performing keyword matching on each piece of data to be analyzed in the at least one piece of data to be analyzed to determine at least one entity label of the data to be analyzed includes: Based on each entity keyword in the first preset rule, keyword matching is performed on the data to be analyzed to determine at least one candidate keyword in the data to be analyzed; and In response to the inclusion of the negative keyword among the at least one candidate keyword, the corresponding keyword among the at least one candidate keyword is filtered out, so as to determine at least one entity label of the data to be analyzed based on the remaining keywords among the at least one candidate keyword.
5. The method according to claim 1 or 2, wherein obtaining the user search text from the at least one set of data to be analyzed includes: Retrieve multiple first-user search texts; For each of the plurality of first user search texts, industry category prediction is performed to determine the industry category to which each first user search text belongs; as well as At least one user search text whose industry category is the target industry category is obtained from the plurality of first user search texts, and is used as at least one user search text in the at least one data to be analyzed.
6. The method according to claim 1 or 2, wherein obtaining the webpage title from the at least one set of data to be analyzed includes: Obtain at least one target site, wherein the content category of each target site corresponds to the target industry category; as well as Extract the page title contained in each page of the at least one target site, and use it as at least one page title in the at least one set of data to be analyzed.
7. A data processing apparatus, the apparatus comprising: The first acquisition unit is configured to acquire at least one piece of data to be analyzed, wherein the data to be analyzed includes at least one of user search text and webpage title; The second acquisition unit is configured to acquire at least one preset rule, wherein each preset rule includes at least one entity keyword; The matching unit is configured to perform keyword matching on each of the at least one pieces of data to be analyzed based on each entity keyword in the at least one preset rule, so as to determine at least one entity tag of the data to be analyzed, wherein each entity tag in the at least one entity tag corresponds to the matched entity keyword; An aggregation unit is configured to aggregate user behavior data associated with each of the at least one data to be analyzed, based on at least one entity label of each data to be analyzed, to obtain a set of user behavior data corresponding to each entity label; and The analysis unit is configured to perform statistical analysis based on the user behavior data set corresponding to each entity tag in order to obtain business data analysis results.
8. The apparatus according to claim 7, further comprising: An execution unit is configured to perform the operations of the following subunits for each of the at least one set of data to be analyzed, the execution unit comprising: The word segmentation subunit is configured to segment the data to be analyzed to obtain multiple words; The identification subunit is configured to perform named entity recognition on the plurality of word segments to obtain the entity type of each word segment in the plurality of word segments; and The annotation subunit is configured to annotate the at least two segmented words as common entities in response to at least two segmented words having the same entity type. The update unit is configured to, in response to a situation where the statistical probability of a first co-occurring entity in the at least one set of data to be analyzed is greater than a preset probability threshold, a first word segment in the first co-occurring entity does not match an entity keyword in the at least one preset rule, and a second word segment in the first co-occurring entity matches an entity keyword in the at least one preset rule, add the first word segment as an entity keyword to the preset rule corresponding to the second word segment to update the at least one preset rule; and The matching unit is further configured to perform keyword matching on each piece of data to be analyzed in the at least one set of data to be analyzed, based on an updated preset rule, so as to update the entity label of the data to be analyzed.
9. The apparatus according to claim 7 or 8, wherein, The number of the at least one preset rule is multiple, and the first rule and the second rule in the at least one preset rule have a hierarchical relationship in terms of entities, and the entity type of the first rule is the parent type of the entity type of the second rule. The matching unit further includes: A subunit is determined, configured to respond to at least one entity tag in the data to be analyzed including an entity tag corresponding to an entity keyword in the second rule, and to determine the entity tag of the data to be analyzed that corresponds to the corresponding entity keyword in the first rule based on the hierarchical relationship between the entities in the first rule and the second rule.
10. The apparatus according to claim 7 or 8, wherein, The first preset rule in the at least one preset rule includes a negative keyword, and the matching unit includes: The matching subunit is configured to perform keyword matching on the data to be analyzed based on each entity keyword in the first preset rule, to determine at least one candidate keyword in the data to be analyzed; and The filtering subunit is configured to filter out corresponding keywords from the at least one candidate keywords in response to the inclusion of the negative keyword in the at least one candidate keywords, so as to determine at least one entity label of the data to be analyzed based on the remaining keywords in the at least one candidate keywords.
11. The apparatus according to claim 7 or 8, wherein the first acquiring unit comprises: The first acquisition subunit is configured to acquire multiple first user search texts; The prediction subunit is configured to predict the industry category for each of the plurality of first user search texts to determine the industry category to which each first user search text belongs; as well as The second acquisition subunit is configured to acquire at least one user search text whose industry category is the target industry category from the plurality of first user search texts, as at least one user search text in the at least one data to be analyzed.
12. The apparatus according to claim 7 or 8, wherein the first acquiring unit comprises: The third acquisition subunit is configured to acquire at least one target site, wherein the content category of each target site corresponds to a target industry category; as well as The extraction subunit is configured to extract the webpage title contained in each webpage of the at least one target site as at least one webpage title in the at least one set of data to be analyzed.
13. An electronic device, comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.
15. A computer program product comprising a computer program, wherein, The computer program, when executed by a processor, implements the method of any one of claims 1-6.
Citation Information
Patent Citations
Teaching problem diagnosis method and system based on knowledge graph
CN110083744A
Data analysis method and device based on search engine
CN111475536A