• [技术干货] 华为算法精英实战营第4期 - 云集群成本优化 参赛经验分享
    赛题核心点介绍这道题的核心赛点在于设计算法将结点分配给适当类型的服务器来最小化费用。整体思路介绍1. 对每一个CREATE操作,把同类(CPU和MEM相同)的查询规为一类,先看看之前有没有空位可填,剩下的做一个最优化dp。对于任意一种类型,令dp[n]表示给n个结点分配服务器的最小代价。则有两种转移,第一种把所有结点塞到一个服务器里,第二种是dp[n] = min{dp[i] + dp[n - i]}。2. 注意每个服务器只存放同类的结点就行了。这样做的原因是同类结点持续时间一般相同,更有可能同时删除,利用率会更高。更具体一点,就比如说在调用new分配内存的时候,同样大小的块是会被放在一起的,因为他们可能会被同时访问,这样基于空间局部性性能就会更高。放到这题里就是:同类结点的持续时间更一致,放到一组里利用率会更高。因为给一个结点分配了服务器之后就不能改变了,而且当服务器空闲的时候会自动释放。3. 一个服务器只能放一个类型结点,如果有结点结束了,之后往里面添加的,也必须同类型。也就是说服务器对应的类型在创建的时候就确定了。核心策略详解当着前面的只是基本思路,要像得到更高的分数需要更多的优化:1. 维护实时的服务器占用率,对每种类型的服务器维护一个占用率的期望值,然后在dp分配结点的时候用 价格/利用率 作为这种服务器的真实价格。对于单台服务器,其占用率的定义为:所有在这个服务器上执行过的结点的持续时长的和 / (服务器存在的时长 * 能存放的结点数量)。对于一个类型的服务器,其占用率的期望定义为到目前为止所有该类型服务器占用率的平均值。2. 基于上面的策略,假设对于某个类型x的结点,其最优的服务器类型是y,那么对于一次CREATE操作,我们的dp算法要做的就是分配足够数量的y服务器,此时剩下的少量结点会分配非y类型的服务器。举例来说:假设y能放10个x的结点,但是如果CREATE了11个,那么我们只会创建1一个y,剩下单出来的x有一个容量更小的服务器运行即可。这个算法看起来很合理,但是如果下一次CREATE操作很快就会到来的话,不如直接一次创建2个y,虽然会有一个y的占用率很低,但是这个低占用率的持续时间是很短的。对此我们要做的事情也就很明确了:根据之前的CREATE操作维护一些信息来预测下一次CREATE到来的期望时间,然后根据这个信息来决定最后一次分配是选择y还是更小型的服务器。
  • [技术干货] 第五期 - 高维向量数据的近似检索 RANK2 团队解题思路分享
    高维向量数据的近似检索-huaweifirst_worldlatter团队赛题核心点介绍1.采用c++编写可执行程序单纯python+numpy的版本分数在1600.0374左右,上限较低,改用纯c++实现可以达到1600.0474分左右2.采用循环展开由于无法开启编译器simd指令优化,因此采用了循环展开(采用4/8/16)展开。对于M个doc向量在与query计算时可以使用循环展开,每次计算N(4/8/16)个doc的相似度,编译器会尝试生成向量化指令,从而加速整体的运算速度3.采用简化数据计算根据(x-y)^2=x^2+y^2-2xy可知,如果将向量doc提前归一化后可以将x^2和y^2的计算舍弃,进而只有一次乘法近似计算4.采用cache对于已经计算过的输入则不会重复计算,没有计算过的输入会保存其输出到cache中方便之后相同输出差表。整体思路介绍核心策略详解1.提前分配内存避免后面动态分配内存造成额外耗时由于已知doc的数量和每次只查询一个query,因此占用的内存是固定,可以提前开辟一块连续的大块内存,避免后面算法在动态分配内存造成内存碎片。2.数据读取阶段加速,关闭io相关的同步,加快数据读取由于c++的流机制存在的同步,因此关闭同步并且解绑输入输出流可以加快io的速度。3.通过使用doc归一化手段简化欧式距离计算,只用一次乘法即可获取结果简化了欧式距离的流程,由原来长度为L的向量需要计算两次减法一次乘法一次跟号共4L次运算,变换为只需要计算一次乘法L次运算。4.通过循环展开由于无法使用simd,因此采用循环展开,对于M个doc进行分块,按照4/8/16分块后每次计算4/8/16个doc与query的欧式距离,最后将剩余的块单独计算。经过编译器优化后可以产生一次加载4/8/16元素的向量化访存和计算的指令,从而加快整个doc的计算。5.通过优先队列计算topk由于topk的计算不需要全部doc的分数进行排序,只需要排序好前topk个即可,因此采用了堆排序的优先队列,可以避免全量doc的分数排序。6.使用cache缓存由于推理时可能有重复的query请求,为了避免重复query请求浪费计算力,会在计算时缓存旧的结果,如果遇到已经缓存的结果则直接取缓存值。
  • [大赛资讯] 【第六期拓扑感知的虚拟机放置问题】赛题补充介绍以及相关的一些算法思路建议。
    附件为出题专家组撰写的赛题介绍,以及关于第六期赛题的一些算法思路建议,对于各位选手的解题相信会有很大的启发,大家赶紧下载查看转发起来吧!以下为正文:Cloud computing has emerged as a new paradigm for on-demand access to a huge pool of computing resources. During the last decade it has formed a large-scale ecosystem of technologies, products and services which enabled the rapid innovation and development of new applications with millions of users. At the same time, cloud computing has posed new challenging problems to cloud providers that should operate the huge resource pools in a cost efficient way while providing a high quality of service to their users.To win the competition, major cloud service providers are focusing on improving the utilization of cloud resources, as this will allow to reduce the cost of cloud services. It is estimated that even a 1% increase in cloud service resource utilization can save millions of dollars in costs. The optimization of resource utilization, in turn, requires the development of advanced algorithms and technologies for cloud resource management. The virtual machine (VM) placement algorithm considered in this contest is one of these critical algorithms. This algorithm works online by finding a suitable server for running each requested user VM. The algorithm should compute its decision fast while providing a dense packing of VMs and avoiding rejections due to the lack of free space. The increase of the scale and complexity of cloud services and deployed applications has led to new user requirements. In particular, instead of a single VM, users often need to deploy a set of VMs with specific requirements on the quality of their placement. For example, to ensure high availability each VM must be placed in a separate fault domain, a set of servers that share a single point of failure, such as rack. In other situation, like neural network training, it may be preferable to place VMs "close" to each other in terms of cloud network topology to reduce the communication latency and speed up the computations.These new requirements bring new challenges to VM placement problem – how to quickly and efficiently place large sets of VMs while satisfying numerous constraints. Basic algorithms that place each VM independently cannot adapt to such complex scenarios. So the context of this Challenge is development of topology-aware VM placement algorithm that can both satisfy new constraints and ensure efficient resource utilization.The algorithm operates a large pool of servers by processing incoming user requests on creation and deletion of virtual machines. The resource pool is organized in a hierarchical topology with multiple levels. Servers live in racks with form separate fault domains, because the failure of the rack network switch or power supply leads to the failure of all servers from the rack. Racks are grouped into pods and network domains. The network communication inside the network domain is much faster than between different network domains.Cloud users create VMs from the predefined set of VM types. A user can request creation of multiple VMs of the same type at once. A user can request the deletion of any previously created VM at any time. Each VM is associated with some placement group (PG). The main role of PG is to define the requirements on placement of VMs belonging to this group. These requirements are specified as affinity and anti-affinity constraints. For example, each VM in group should be placed in a separate rack is an anti-affinity constraint. It is also possible to combine multiple constraints in a PG.The algorithm processes requests one by one, in online fashion. That means that the algorithm must output its decision regarding the current request without knowledge about the next requests. This is exactly how such algorithm runs in a real cloud. Your main goal will be to maximize the number of successfully placed VMs. As the algorithm runs on a resource pool of fixed size the more requests are allocated, the less resources are wasted and the more resource efficient is the algorithm.The first version of this problem was published as the ICPC 2022 Online Challenge powered by Huawei which attracted many algorithm enthusiasts from all over the world. For the new competition, the problem has been reworked to introduce more complex topology and placement constraints, new test cases and solution scoring. Below we describe the main differences.To model new placement constraints arising in AI training scenarios, a new topology level (pod) has been introduced which corresponds to a set of racks. Training of large AI models requires a massive amount of computations and typically runs in parallel across a large number of servers. The communication speed between these servers is crucial, therefore such workloads require allocation of one or several whole pods in a single network domain. Such constraints has been added to this challenge. We have also added more complex real-world constraints which combine affinity and anti-affinity on different topology levels.57 completely new test cases have been prepared for this challenge. While the previous challenge featured only synthetic test cases, this challenge also features test cases based on real cloud requests. The generator of synthetics cases has also been improved.Finally, the solution scoring is also reworked. While in ICPC Challenge the achieved results were compared with simple baseline algorithm, now we compare them with an upper bound on the number of successfully placed VMs which is a hard theoretical limit.The problem being solved shares many similarities with the classic bin packing and knapsack problems. Indeed, we need to pack items represented by VMs into bins represented by servers. Therefore algorithms for solving these well-known problems can be used to efficiently pack VMs into servers without leaving unused resources. However, our problem is remarkably harder than these classic problems because of the following features. First, we consider online scheduling in which future requests are unknown in advance, while most of classic algorithms consider offline setting where all packed items are given as input. Development of online algorithm requires dealing with uncertainty. One approach is to minimize the risks by preparing to challenging requests. For example, we can be prepared for upcoming requests for large VM by reserving enough free space to accommodate such VMs. Another approach is to try to anticipate future requests by using some predictive model based on already processed requests. For example, the distribution of requested VM types, or what kind or requests arrived recently, may provide some hints about the upcoming requests.Second, we consider the dynamic packing problem where the items may depart from the bins, which also quite different from the classic bin packing. While item departures are not as challenging as dynamic item arrivals and can even simplify the packing, they can also be leveraged to improve the packing efficiency. For example, by packing together VMs with similar lifetime, we can improve the availability of large free space on servers for packing large VMs. This requires predicting the VM lifetime, which can be also done using the already observed information. Note also that VMs from the same PG are usually deleted in batches, so their lifetime is similar.Third, the most complexity comes from the placement constraints. Consider a PG with affinity constraint on rack level. When placing the first request from this group, we need to choose a rack with enough free space to accommodate the VMs from this request and the VMs from upcoming requests of this group. This means that it is no longer safe to pack VMs one by one, without considering whether the whole request will fit a rack. It also means that a care should be taken to provide some room for possible future growth of the group inside the chosen rack. One possible approach is to choose a rack with the most available free space. Note also that groups cannot grow infinitely, there is some limit on their size. Anti-affinity constraints are usually less challenging but they also require the strategical management of available space to success. For example, for anti-affinity on rack level with partitions it is better to choose the racks where the request partition is already present to avoid occupying too much racks for some partition, since it will reduce possible racks for accommodating other partitions. Think about what kind of similar strategies can be applied for other constraints. Finally, when choosing a topology domain for placing VMs it is convenient to select domains in a top-down fashion using some scoring metrics or rules for each domain on some level.As a last recommendation and a word of caution, try to avoid overfitting your solution parameters and even placement strategies to individual tests. This will both complicate your algorithm and make it inapplicable in a real cloud setting. Strive for simpler and general algorithms with a clear intuition behind them and a small number of parameters, and also add comments in the code to explaining them. That is what we expect from a perfect solution to our challenge.
  • [第5期高维向量数据] total time limit: 50 seconds是指每个测试集50秒还是一共50秒
    total time limit: 50 seconds是指每个测试集50秒还是一共50秒
  • [第5期高维向量数据] 得分为0,反馈信息也没有,
    但是本地测试一些简单的样例后,结果确实是正确的。baseline有分数,但是使用一些近似检索算法后就一直得分为0。。图1是baseline,图2 是改进代码,没有调用额外的库。
  • [常见FAQ] python程序在上传之后显示”初始化超时“
    在本地运行没有任何问题,本次初始化5ms都不到的初始化程序,在提交到系统之后一直显示”初始化超时“,求指导,该怎么办?
  • [应用开发] SegFormer-B0 OM模型在MDC300F上推理时间为1000ms
    在将segformer-b0算法模型转OM后,在MDC300F MINI上推理耗时1000ms。ONNX中MatMul算子的输入数据的shape带batch维度,将算子的type类型更改为 BatchMatMul后,转OM时,发现MatMul算子前后各出现一个trans_TransData算子。通过profiling计算算子耗时,发现TransposeD、ArgMaxD、TransData、SoftmaxV2、BatchMatMul占用大量的推理时间。请问如何避免带batch维度的MatMul产生TransData算子?针对现在算子耗时,有什么优化的方法吗?
  • [技术干货] 深度优先搜索(DFS)与广度优先搜索(BFS)及其应用场景
    在算法领域中,深度优先搜索(DFS)和广度优先搜索(BFS)是两种常见且重要的图遍历算法。它们各自具有独特的特点和适用场景。本文将详细介绍这两种算法,并探讨它们的最佳应用场景。一、深度优先搜索(DFS)深度优先搜索是一种用于遍历或搜索树或图的算法。这个算法会尽可能深地搜索图的分支。当节点v的所在边都己被探寻过,搜索将回溯到发现节点v的那条边的起始节点。这一过程一直进行到已发现从源节点可达的所有节点为止。如果还存在未被发现的节点,则选择其中一个作为源节点并重复以上过程,整个进程反复进行直到所有节点都被访问为止。DFS的应用场景:图的连通性检查:给定一个无向图,判断是否存在从节点A到节点B的路径。 生成树的构造:在计算机网络中,DFS可以用来发现网络的拓扑结构,并生成一棵生成树。 解决迷宫问题:在迷宫问题中,DFS可以用来找到从入口到出口的路径。 树的遍历:对于二叉树、多叉树等数据结构,DFS是一种非常有效的遍历方式。二、广度优先搜索(BFS)广度优先搜索是另一种用于遍历或搜索树或图的算法。它从根(或任意节点)开始并探索最近的节点,然后进行下一层的邻居节点,这样一层一层地进行。BFS使用队列数据结构来保存信息。BFS的应用场景:最短路径问题:在图中找到从一个节点到另一个节点的最短路径。例如,在地图中找到从起点到终点的最短路线。 层次遍历:对于二叉树、多叉树等数据结构,BFS可以用来进行层次遍历。 图的连通性检查:与DFS相似,BFS也可以用来检查图的连通性,但它通常用于找到从起点到所有其他可达节点的最短路径。 网络爬虫:在网页搜索中,BFS被用来遍历网页之间的链接,从而找到与给定网页相关的所有网页。总结DFS和BFS各有优缺点,选择哪种算法取决于具体的问题和应用场景。DFS通常用于深度探索,而BFS则更适用于广度探索。在实际应用中,我们需要根据问题的特性来选择最合适的算法。希望本文能帮助您更好地理解深度优先搜索和广度优先搜索,以及它们在不同场景下的应用。
  • [大赛资讯] 关于A题
    节点被自动删除后,新生成的节点编号还会复用此前的吗?比如样例中3号节点在4号pod被删除后就删除了,那再新生成一个节点编号是3还是4?
  • [技术干货] 开源项目分享之Hello算法(hello-algo)
    动画图解、一键运行的数据结构与算法教程项目地址: GitHub在线阅读                                   下载PDF《Hello 算法》:动画图解、一键运行的数据结构与算法教程,支持 Java, C++, Python, Go, JS, TS, C#, Swift, Rust, Dart, Zig 等语言。​关于本书本项目旨在打造一本开源免费、新手友好的数据结构与算法入门教程。全书采用动画图解,内容清晰易懂、学习曲线平滑,引导初学者探索数据结构与算法的知识地图。源代码可一键运行,帮助读者在练习中提升编程技能,了解算法工作原理和数据结构底层实现。鼓励读者互助学习,提问与评论通常可在两日内得到回复。推荐语
  • [互动交流] 雪花算法生成的ID为什么能保证全局唯一?
    雪花算法生成的ID为什么能保证全局唯一?
  • [其他] 浅谈AdaBoost算法
     Adaboost是一种迭代算法,其核心思想是针对同一个训练集训练不同的分类器(弱分类器),然后把这些弱分类器集合起来,构成一个更强的最终分类器(强分类器)。其算法本身是通过改变数据分布来实现的,它根据每次训练集之中每个样本的分类是否正确,以及上次的总体分类的准确率,来确定每个样本的权值。将修改过权值的新数据集送给下层分类器进行训练,最后将每次训练得到的分类器最后融合起来,作为最后的决策分类器。使用adaboost分类器可以排除一些不必要的训练数据特征,并放在关键的训练数据上面。 AdaBoost用于短决策树。创建第一棵树之后,使用树在每个训练实例上的性能来得到一个权重,决定下一棵树对每个训练实例的注意力。 难以预测的训练数据被赋予更多权重,而易于预测的实例被赋予更少的权重。模型是一个接一个地顺序创建的,每个模型更新训练实例上的权重,这些权重影响序列中下一个树所执行的学习。构建完所有树之后,将对新数据进行预测。 因为着重于修正算法的错误,所以重要的是提前清洗好数据,去掉异常值。 Adaboost是一种迭代算法,其核心思想是针对同一个训练集训练不同的分类器(弱分类器),然后把这些弱分类器集合起来,构成一个更强的最终分类器(强分类器)。  算法应用  对adaBoost算法的研究以及应用大多集中于分类问题,同时也出现了一些在回归问题上的应用。就其应用adaBoost系列主要解决了: 两类问题、多类单标签问题、多类多标签问题、大类单标签问题、回归问题。它用全部的训练样本进行学习。 应用示例  adaboost 是 bosting 的方法之一,bosting就是把若干个分类效果并不好的分类器综合起来考虑,会得到一个效果比较好的分类器。 下图,左右两个决策树,单个看是效果不怎么好的,但是把同样的数据投入进去,把两个结果加起来考虑,就会增加可信度。  adaboost 的栗子,手写识别中,在画板上可以抓取到很多 features,例如 始点的方向,始点和终点的距离等等。  training 的时候,会得到每个 feature 的 weight,例如 2 和 3 的开头部分很像,这个 feature 对分类起到的作用很小,它的权重也就会较小。  而这个 alpha 角 就具有很强的识别性,这个 feature 的权重就会较大,最后的预测结果是综合考虑这些 feature 的结果。 
  • [技术干货] wrf非压缩版安装
    wrf压缩版和非压缩版的区别?wrf运行的输出文件是否被压缩wrf压缩版和非压缩版的关键压缩wrf的中输出文件由netcdf处理输出,所以两者区别在于netcdf的不同。社区文档默认安装压缩版的netcdf,现在记录下非压缩版netcdf的安装压缩版netcdf-c,主要添加--disable-netcdf-4./configure --prefix=${NETCDT_DIR} --build=aarch64-unknown-linux-gnu --disable-netcdf-4 --enable-shared --enable-netcdf-4 --disable-dap --with-pic --disable-doxygen --enable-static --enable-pnetcdf --enable-largefile CPPFLAGS="-O3 -I${HDF5_DIR}/include -I${PNETCDF_DIR}/include" LDFLAGS="-L${HDF5_DIR}/lib -L${PNETCDF_DIR}/lib -Wl,-rpath=${HDF5_DIR}/lib -Wl,-rpath=${PNETCDF_DIR}/lib" CFLAGS="-O3 -L${HDF5_DIR}/lib -L${PNETCDF_DIR}/lib -I${HDF5_DIR}/include -I${PNETCDF_DIR}/include"压缩版netcdf-fortran,主要添加--disable-fortran-type-check./configure --prefix=${NETCDT_DIR} --build=aarch64-unknown-linux-gnu --disable-fortran-type-check --enable-shared --with-pic --disable-doxygen --enable-largefile --enable-static CPPFLAGS="-O3 -I${HDF5_DIR}/include -I${NETCDT_DIR}/include" LDFLAGS="-L${HDF5_DIR}/lib -L${NETCDT_DIR}/lib -Wl,-rpath=${HDF5_DIR}/lib -Wl,-rpath=${NETCDT_DIR}/lib" CFLAGS="-O3 -L${HDF5_DIR}/lib -L${NETCDT_DIR}/lib -I${HDF5_DIR}/include -I${NETCDT_DIR}/include" CXXFLAGS="-O3 -L${HDF5_DIR}/lib -L${NETCDT_DIR}/lib -I${HDF5_DIR}/include -I${NETCDT_DIR}/include" FCFLAGS="-O3 -L${HDF5_DIR}/lib -L${NETCDT_DIR}/lib -I${HDF5_DIR}/include -I${NETCDT_DIR}/include"其余安装步骤与压缩版netcdf移植。检查netcdf是否是压缩版命令行输入: nc-config --all 非压缩版如下图:压缩版如下图:判断wrf是否是压缩版:./configure 之后出现以下信息: 非压缩版: 压缩版:
  • [技术干货] One-Class SVM介绍【转】
    传统意义上,很多的分类问题试图解决两类或者多类情况,机器学习应用的目标是采用训练数据,将测试数据属于哪个类进行区分。但是如果只有一类数据,而目标是测试新的数据并且检测它是否与训练数据相似。One-Class Support Vector Machine在过去20里逐渐流行,就是为了解决这样的问题。本篇博客对One-Class SVM进行介绍,而One-Class SVM也会应用到我的论文中。1. Just One Class?首先先看下我们的问题,我们希望确定新的训练数据是否属于某个特定类?怎么会有这样的应用场景呢?比方说,我们要判断一张照片里的人脸,是男性还是女性,这是个二分类问题。就是将一堆已标注了男女性别的人脸照片(假设男性是正样本,女性是负样本),提取出有区分性别的特征(假设这种能区分男女性别的特征已构建好)后,通过svm中的支持向量,找到这男女两类性别特征点的最大间隔。进而在输入一张未知性别的照片后,经过特征提取步骤,就可以通过这个训练好的svm很快得出照片内人物的性别,此时我们得出的结论,我们知道要么是男性,不是男性的话,那一定是女性。以上情况是假设,我们用于训练的样本,包括了男女两类的图片,并且两类图片的数目较为均衡。现实生活中的我们也是这样,我们只有在接触了足够多的男生女生,知道了男生女生的性别特征差异后(比方说女性一般有长头发,男性一般有胡子等等),才能准确判断出这个人到底是男是女。但如果有一个特殊的场景,比方说,有一个小和尚,他从小在寺庙长大,从来都只见过男生,没见过女生,那么对于一个男性,他能很快地基于这个男性与他之前接触的男性有类似的特征,给出这个人是男性的答案。但如果是一个女性,他会发现,这个女性与他之前所认知的男性的特征差异很大,进而会得出她不是男性的判断。但注意这里的“她不是男性”不能直接归结为她是女性。再比如,在工厂设备中,控制系统的任务是判断是是否有意外情况出现,例如产品质量过低、机器产生奇怪的震动或者温度上升等。相对来说容易收集到正常场景下的训练数据,但故障系统状态的收集示例数据可能相当昂贵,或者根本不可能。 如果可以模拟一个错误的系统状态,但必然无法保证所有的错误状态都被模拟到,从而在传统的两类问题中得到识别。2. Schölkopf的OCSVMThe Support Vector Method For Novelty Detection by Schölkopf et al. 基本上将所有的数据点与零点在特征空间分离开,并且最大化分离超平面到零点的距离。这产生一个binary函数能够获取特征空间中数据的概率密度区域。当处于训练数据点区域时,返回+1,处于其他区域返回-1.该问题的优化目标与二分类SVM略微不同,但依然很相似其中表示松弛变量, 类似于二分类SVM中的,同时:它为异常值的分数设置了一个上限(训练数据集里面被认为是异常的)是训练数据集里面做为支持向量的样例数量的下届因为这个参数的重要性,这种方法也被称为。采用Lagrange技术并且采用dot-product calculation,确定函数变为:这个方法创建了一个参数为的超平面,该超平面与特征空间中的零点距离最大,并且将零点与所有的数据点分隔开。3. SVDDThe method of Support Vector Data Description by Tax and Duin (SVDD)采用一个球形而不是平面的方法,该算法在特征空间中获得数据周围的球形边界,这个超球体的体积是最小化的,从而最小化异常点的影响。产生的超球体参数为中心和半径 ,体积 被最小化,中心 是支持向量的线性组合;跟传统SVM方法相似,可以要求所有数据点 到中心的距离严格小于 ,但同时构造一个惩罚系数为 的松弛变量,优化问题如下所示:在采用拉格朗日算子求解之后,可以判断新的数据点是否在类内,如果z到中心的距离小于或者等于半径。采用Gaussian Kernel做为两个数据点的距离函数:转自 https://zhuanlan.zhihu.com/p/32784067
  • [技术干货] 情感分析神器!再也不怕女朋友生气 【转】
    背景介绍众所周知,人类自然语言中包含了丰富的情感色彩:表达人的情绪(如悲伤、快乐)、表达人的心情(如倦怠、忧郁)、表达人的喜好(如喜欢、讨厌)、表达人的个性特征和表达人的立场等等。情感分析在商品喜好、消费决策、舆情分析等场景中均有应用。利用机器自动分析这些情感倾向,不但有助于帮助企业了解消费者对其产品的感受,为产品改进提供依据;同时还有助于企业分析商业伙伴们的态度,以便更好地进行商业决策。被人们所熟知的情感分析任务是将一段文本分类,如分为情感极性为 正向、负向、其他的三分类问题:情感分析任务正向:表示正面积极的情感,如高兴,幸福,惊喜,期待等。负向:表示负面消极的情感,如难过,伤心,愤怒,惊恐等。其他:其他类型的情感。实际上,以上熟悉的情感分析任务是句子级情感分析任务。情感分析预训练模型SKEP近年来,大量的研究表明基于大型语料库的预训练模型(Pretrained Models, PTM)可以学习通用的语言表示,有利于下游NLP任务,同时能够避免从零开始训练模型。随着计算能力的发展,深度模型的出现(即 Transformer)和训练技巧的增强使得 PTM 不断发展,由浅变深。情感预训练模型SKEP(Sentiment Knowledge Enhanced Pre-training for Sentiment Analysis)。SKEP利用情感知识增强预训练模型, 在14项中英情感分析典型任务上全面超越SOTA,此工作已经被ACL 2020录用。SKEP是百度研究团队提出的基于情感知识增强的情感预训练算法,此算法采用无监督方法自动挖掘情感知识,然后利用情感知识构建预训练目标,从而让机器学会理解情感语义。SKEP为各类情感分析任务提供统一且强大的情感语义表示。论文地址:https://arxiv.org/abs/2005.05635百度研究团队在三个典型情感分析任务,句子级情感分类(Sentence-level Sentiment Classification),评价目标级情感分类(Aspect-level Sentiment Classification)、观点抽取(Opinion Role Labeling),共计14个中英文数据上进一步验证了情感预训练模型SKEP的效果。具体实验效果参考:https://github.com/baidu/Senta#skep情感分析任务还可以进一步分为句子级情感分析、目标级情感分析等任务。在下面章节将会详细介绍两种任务及其应用场景。句子级情感分析 &目标级情感分析本项目将详细全面介绍情感分析任务的两种子任务,句子级情感分析和目标级情感分析。记得给 PaddleNLP点个小小的 Star⭐开源不易,希望大家多多支持~GitHub地址:https://github.com/PaddlePaddle/PaddleNLPAI Studio平台后续会默认安装PaddleNLP最新版,在此之前可使用如下命令更新安装。!pip install--upgrade paddlenlp -i https://pypi.org/simple1 句子级情感分析对给定的一段文本进行情感极性分类,常用于影评分析、网络论坛舆情分析等场景。如:① 选择珠江花园的原因就是方便,有电动扶梯直接到达海边,周围餐馆、食廊、商场、超市、摊位一应俱全。酒店装修一般,但还算整洁。泳池在大堂的屋顶,因此很小,不过女儿倒是喜欢。包的早餐是西式的,还算丰富。服务吗,一般 1② 1 5.4寸笔记本的键盘确实爽,基本跟台式机差不多了,蛮喜欢数字小键盘,输数字特方便,样子也很美观,做工也相当不错 1③ 房间太小。其他的都一般。。。。。。。。。 0其中1表示正向情感,0表示负向情感。句子级情感分析任务常用数据集ChnSenticorp数据集是公开中文情感分析常用数据集, 其为二分类数据集。PaddleNLP已经内置该数据集,一键即可加载。frompaddlenlp.datasets importload_datasettrain_ds, dev_ds, test_ds = load_dataset( "chnsenticorp", splits=[ "train", "dev", "test"])print(train_ds[ 0])print(train_ds[ 1])print(train_ds[ 2]){ 'text': '选择珠江花园的原因就是方便,有电动扶梯直接到达海边,周围餐馆、食廊、商场、超市、摊位一应俱全。酒店装修一般,但还算整洁。 泳池在大堂的屋顶,因此很小,不过女儿倒是喜欢。 包的早餐是西式的,还算丰富。 服务吗,一般', 'label': 1, 'qid': ''}1.1 SKEP模型加载PaddleNLP已经实现了SKEP预训练模型,可以通过一行代码实现SKEP加载。句子级情感分析模型是SKEP fine-tune 文本分类常用模型SkepForSequenceClassification。其首先通过SKEP提取句子语义特征,之后将语义特征进行分类。from paddlenlp.transformers import SkepForSequenceClassification, SkepTokenizer# 指定模型名称,一键加载模型model = SkepForSequenceClassification.from_pretrained(pretrained_model_name_or_path= "skep_ernie_1.0_large_ch", num_classes=len(train_ds.label_list)) # 同样地,通过指定模型名称一键加载对应的Tokenizer,用于处理文本数据,如切分token,转token_id等。tokenizer = SkepTokenizer.from_pretrained(pretrained_model_name_or_path= "skep_ernie_1.0_large_ch")SkepForSequenceClassification可用于句子级情感分析和目标级情感分析任务。其通过预训练模型SKEP获取输入文本的表示,之后将文本表示进行分类。pretrained_model_name_or_path:模型名称。支持"skep_ernie_1.0_large_ch","skep_ernie_2.0_large_en"。"skep_ernie_1.0_large_ch":是SKEP模型在预训练ernie_1.0_large_ch基础之上在海量中文数据上继续预训练得到的中文预训练模型;"skep_ernie_2.0_large_en":是SKEP模型在预训练ernie_2.0_large_en基础之上在海量英文数据上继续预训练得到的英文预训练模型;num_classes: 数据集分类类别数。关于SKEP模型实现详细信息参考:https://github.com/PaddlePaddle/PaddleNLP/tree/develop/paddlenlp/transformers/skep1.2 数据处理同样地,我们需要将原始ChnSentiCorp数据处理成模型可以读入的数据格式。SKEP模型对中文文本处理按照字粒度进行处理,我们可以使用PaddleNLP内置的SkepTokenizer完成一键式处理。importosfromfunctools importpartialimportnumpy asnpimportpaddleimportpaddle.nn.functional asFfrompaddlenlp.data importStack, Tuple, Padfromutils importcreate_dataloaderdefconvert_example(example,tokenizer,max_seq_length= 512,is_test=False) :# 将原数据处理成model可读入的格式,enocded_inputs是一个dict,包含input_ids、token_type_ids等字段encoded_inputs = tokenizer(text=example[ "text"], max_seq_len=max_seq_length)# input_ids:对文本切分token后,在词汇表中对应的token idinput_ids = encoded_inputs[ "input_ids"]# token_type_ids:当前token属于句子1还是句子2,即上述图中表达的segment idstoken_type_ids = encoded_inputs[ "token_type_ids"]ifnotis_test:# label:情感极性类别label = np.array([example[ "label"]], dtype= "int64")returninput_ids, token_type_ids, labelelse:# qid:每条数据的编号qid = np.array([example[ "qid"]], dtype= "int64")returninput_ids, token_type_ids, qid# 批量数据大小batch_size = 32 # 文本序列最大长度max_seq_length = 256# 将数据处理成模型可读入的数据格式trans_func = partial(convert_example,tokenizer=tokenizer,max_seq_length=max_seq_length)# 将数据组成批量式数据,如# 将不同长度的文本序列padding到批量式数据中最大长度# 将每条数据label堆叠在一起batchify_fn = lambda samples, fn=Tuple(Pad(axis=0, pad_val=tokenizer.pad_token_id), # input_idsPad(axis=0, pad_val=tokenizer.pad_token_type_id), # token_type_idsStack # labels): [data for data in fn(samples)]train_data_loader = create_dataloader(train_ds,mode='train',batch_size=batch_size,batchify_fn=batchify_fn,trans_fn=trans_func)dev_data_loader = create_dataloader(dev_ds,mode='dev',batch_size=batch_size,batchify_fn=batchify_fn,trans_fn=trans_func)1.3 模型训练和评估定义损失函数、优化器以及评价指标后,即可开始训练。推荐超参设置:max_seq_length=256batch_size=48learning_rate=2e-5epochs=10实际运行时可以根据显存大小调整batch_size和max_seq_length大小。importtimefromutils importevaluate# 训练轮次epochs = 1# 训练过程中保存模型参数的文件夹ckpt_dir = "skep_ckpt"# len(train_data_loader)一轮训练所需要的step数num_training_steps = len(train_data_loader) * epochs# Adam优化器optimizer = paddle.optimizer.AdamW(learning_rate= 2e-5,parameters=model.parameters) # 交叉熵损失函数criterion = paddle.nn.loss.CrossEntropyLoss # accuracy评价指标metric = paddle.metric.Accuracy# 开启训练global_step = 0tic_train = time.time forepoch inrange( 1, epochs + 1):forstep, batch inenumerate(train_data_loader, start= 1):input_ids, token_type_ids, labels = batch# 喂数据给modellogits = model(input_ids, token_type_ids)# 计算损失函数值loss = criterion(logits, labels)# 预测分类概率值probs = F.softmax(logits, axis= 1)# 计算acccorrect = metric.compute(probs, labels)metric.update(correct)acc = metric.accumulateglobal_step += 1ifglobal_step % 10== 0:print("global step %d, epoch: %d, batch: %d, loss: %.5f, accu: %.5f, speed: %.2f step/s"% (global_step, epoch, step, loss, acc,10/ (time.time - tic_train)))tic_train = time.time# 反向梯度回传,更新参数loss.backwardoptimizer.stepoptimizer.clear_gradifglobal_step % 100== 0:save_dir = os.path.join(ckpt_dir, "model_%d"% global_step)ifnotos.path.exists(save_dir):os.makedirs(save_dir)# 评估当前训练的模型evaluate(model, criterion, metric, dev_data_loader)# 保存当前模型参数等model.save_pretrained(save_dir)# 保存tokenizer的词表等tokenizer.save_pretrained(save_dir)globalstep 10, epoch: 1, batch: 10, loss: 0.59440, accu: 0.60625, speed: 0.76step/sglobalstep 20, epoch: 1, batch: 20, loss: 0.42311, accu: 0.72969, speed: 0.78step/sglobalstep 30, epoch: 1, batch: 30, loss: 0.09735, accu: 0.78750, speed: 0.72step/s··········globalstep 290, epoch: 1, batch: 290, loss: 0.14512, accu: 0.92778, speed: 0.71step/sglobalstep 300, epoch: 1, batch: 300, loss: 0.12470, accu: 0.92781, speed: 0.77step/seval loss: 0.19412, accu: 0.926671.4预测提交结果使用训练得到的模型还可以对文本进行情感预测。importnumpy asnpimportpaddle# 处理测试集数据trans_func = partial(convert_example,tokenizer=tokenizer,max_seq_length=max_seq_length,is_test= True)batchify_fn = lambdasamples, fn=Tuple(Pad(axis= 0, pad_val=tokenizer.pad_token_id), # inputPad(axis= 0, pad_val=tokenizer.pad_token_type_id), # segmentStack # qid): [data fordata infn(samples)]test_data_loader = create_dataloader(test_ds,mode= 'test',batch_size=batch_size,batchify_fn=batchify_fn,trans_fn=trans_func)# 根据实际运行情况,更换加载的参数路径params_path = 'skep_ckp/model_500/model_state.pdparams'ifparams_path andos.path.isfile(params_path):# 加载模型参数state_dict = paddle.load(params_path)model.set_dict(state_dict)print( "Loaded parameters from %s"% params_path)label_map = { 0: '0', 1: '1'}results = [] # 切换model模型为评估模式,关闭dropout等随机因素model.evalforbatch intest_data_loader:input_ids, token_type_ids, qids = batch# 喂数据给模型logits = model(input_ids, token_type_ids)# 预测分类probs = F.softmax(logits, axis= -1)idx = paddle.argmax(probs, axis= 1).numpyidx = idx.tolistlabels = [label_map[i] fori inidx]qids = qids.numpy.tolistresults.extend(zip(qids, labels))res_dir = "./results"ifnotos.path.exists(res_dir):os.makedirs(res_dir) # 写入预测结果with open(os.path.join(res_dir, "ChnSentiCorp.tsv"), 'w', encoding="utf8") as f:f.write( "index\tprediction\n")forqid, label inresults:f.write(str(qid[ 0])+ "\t"+label+ "\n")2 目标级情感分析在电商产品分析场景下,除了分析整体商品的情感极性外,还细化到以商品具体的“方面”为分析主体进行情感分析(aspect-level),如下:这个薯片口味有点咸,太辣了,不过口感很脆。关于薯片的口味方面是一个负向评价(咸,太辣),然而对于口感方面却是一个正向评价(很脆)。我很喜欢夏威夷,就是这边的海鲜太贵了。关于夏威夷是一个正向评价(喜欢),然而对于夏威夷的海鲜却是一个负向评价(价格太贵)。目标级情感分析任务常用数据集千言数据集已提供了许多任务常用数据集。其中情感分析数据集下载链接:https://aistudio.baidu.com/aistudio/competition/detail/50/?isFromLUGE=TRUESE-ABSA16_PHNS数据集是关于手机的目标级情感分析数据集。PaddleNLP已经内置了该数据集,加载方式,如下:train_ds, test_ds = load_dataset( "seabsa16", "phns", splits=[ "train", "test"])2.1 SKEP模型加载目标级情感分析模型同样使用SkepForSequenceClassification模型,但目标级情感分析模型的输入不单单是一个句子,而是句对。一个句子描述“评价对象方面(aspect)”,另一个句子描述"对该方面的评论"。如下图所示。# 指定模型名称一键加载模型model = SkepForSequenceClassification.from_pretrained('skep_ernie_1.0_large_ch', num_classes=len(train_ds.label_list)) # 指定模型名称一键加载tokenizertokenizer = SkepTokenizer.from_pretrained( 'skep_ernie_1.0_large_ch')2.2 数据处理同样地,我们需要将原始SE_ABSA16_PHNS数据处理成模型可以读入的数据格式。SKEP模型对中文文本处理按照字粒度进行处理,我们可以使用PaddleNLP内置的SkepTokenizer完成一键式处理。fromfunctools importpartialimportosimport timeimportnumpy asnpimportpaddleimportpaddle.nn.functional asFfrompaddlenlp.data importStack, Tuple, Paddefconvert_example(example,tokenizer,max_seq_length= 512,is_test=False,dataset_name= "chnsenticorp") :encoded_inputs = tokenizer(text=example[ "text"],text_pair=example[ "text_pair"],max_seq_len=max_seq_length)input_ids = encoded_inputs[ "input_ids"]token_type_ids = encoded_inputs[ "token_type_ids"]ifnotis_test:label = np.array([example[ "label"]], dtype= "int64")returninput_ids, token_type_ids, labelelse:returninput_ids, token_type_ids# 处理的最大文本序列长度max_seq_length= 256# 批量数据大小batch_size= 16# 将数据处理成model可读入的数据格式trans_func = partial(convert_example,tokenizer=tokenizer,max_seq_length=max_seq_length) # 将数据组成批量式数据,如# 将不同长度的文本序列padding到批量式数据中最大长度# 将每条数据label堆叠在一起batchify_fn = lambdasamples, fn=Tuple(Pad(axis= 0, pad_val=tokenizer.pad_token_id), # input_idsPad(axis= 0, pad_val=tokenizer.pad_token_type_id), # token_type_idsStack(dtype= "int64") # labels): [data fordata infn(samples)]train_data_loader = create_dataloader(train_ds,mode= 'train',batch_size=batch_size,batchify_fn=batchify_fn,trans_fn=trans_func)2.3 模型训练定义损失函数、优化器以及评价指标后,即可开始训练。# 训练轮次epochs = 3# 总共需要训练的step数num_training_steps = len(train_data_loader) * epochs # 优化器optimizer = paddle.optimizer.AdamW(learning_rate= 5e-5,parameters=model.parameters) # 交叉熵损失criterion = paddle.nn.loss.CrossEntropyLoss # Accuracy评价指标metric = paddle.metric.Accuracy# 开启训练ckpt_dir = "skep_aspect"global_step = 0tic_train = time.time forepoch inrange( 1, epochs + 1):forstep, batch inenumerate(train_data_loader, start= 1):input_ids, token_type_ids, labels = batch# 喂数据给modellogits = model(input_ids, token_type_ids)# 计算损失函数值loss = criterion(logits, labels)# 预测分类概率probs = F.softmax(logits, axis= 1)# 计算acccorrect = metric.compute(probs, labels)metric.update(correct)acc = metric.accumulateglobal_step += 1ifglobal_step % 10== 0:print("global step %d, epoch: %d, batch: %d, loss: %.5f, acc: %.5f, speed: %.2f step/s"% (global_step, epoch, step, loss, acc,10/ (time.time - tic_train)))tic_train = time.time# 反向梯度回传,更新参数loss.backwardoptimizer.stepoptimizer.clear_gradifglobal_step % 100== 0:save_dir = os.path.join(ckpt_dir, "model_%d"% global_step)ifnotos.path.exists(save_dir):os.makedirs(save_dir)# 保存模型参数model.save_pretrained(save_dir)# 保存tokenizer的词表等tokenizer.save_pretrained(save_dir)2.4 模型预测使用训练得到的模型还可以对评价对象进行情感预测。@paddle.no_graddef predict(model, data_loader, label_map):model.evalresults = []forbatch indata_loader:input_ids, token_type_ids = batchlogits = model(input_ids, token_type_ids)probs = F.softmax(logits, axis= 1)idx = paddle.argmax(probs, axis= 1).numpyidx = idx.tolistlabels = [label_map[i] fori inidx]results.extend(labels)returnresults# 处理测试集数据label_map = { 0: '0', 1: '1'}trans_func = partial(convert_example,tokenizer=tokenizer,max_seq_length=max_seq_length,is_test= True)batchify_fn = lambdasamples, fn=Tuple(Pad(axis= 0, pad_val=tokenizer.pad_token_id), # input_idsPad(axis= 0, pad_val=tokenizer.pad_token_type_id), # token_type_ids): [data fordata infn(samples)]test_data_loader = create_dataloader(test_ds,mode= 'test',batch_size=batch_size,batchify_fn=batchify_fn,trans_fn=trans_func)# 根据实际运行情况,更换加载的参数路径params_path = 'skep_ckpt/model_900/model_state.pdparams'ifparams_path andos.path.isfile(params_path):# 加载模型参数state_dict = paddle.load(params_path)model.set_dict(state_dict)print( "Loaded parameters from %s"% params_path)results = predict(model, test_data_loader, label_map)动手试一试是不是觉得很有趣呀。小编强烈建议初学者参考上面的代码亲手敲一遍,因为只有这样,才能加深你对代码的理解呦。本次项目对应的代码:https://aistudio.baidu.com/aistudio/projectdetail/1968542除此之外, PaddleNLP提供了多种预训练模型,可一键调用,来更换一下预训练试试吧:https://paddlenlp.readthedocs.io/zh/latest/model_zoo/transformers.html更多PaddleNLP信息,欢迎访问GitHub点star收藏后体验:https://github.com/PaddlePaddle/PaddleNLP转自 https://www.sohu.com/a/476337708_827544
总条数:211 到第 页
上滑加载中