I am very new to spark and try to develop it on Windows using IntelliJ. This kind environment is not typical environment for developing spark because normally people use Ubuntu + IntelliJ.
I copied a famous example from spark official example SparkPi, when run it in idea, it pops an error:
I did a lot of search but no luck there. Finally I find out that the problem is caused by my pom.xml settings.
org.apache.sparkspark-core_2.101.4.1${myscope}
The reason of doing this is because we don't to package spark-core jar file in our finally package, but when specify it as provided it cannot run locally. The details can be check in another post of mine: IntelliJ "Provided" Scope Problem
The solution is that:
1. keep the ${myscope} value here
2. add pom.xml a variable called myscope
compile
3. In maven part, keep the setting
clean install -Dmyscope=provided
Now the spark object can run successfully and also package as we expected.
If you think this article is useful, please click the ads on this page to help. Thank you very much.
GitHub is a very famous tool but I am just starting to use it... Shy... Because we come from ancient times in source codes management: ClearCase, SVN, (TeamForge), now it's time to embrace GitHub.
I have downloaded a windows version. GitHub for Windows... Maybe the best way is to use the command line which I will investigate later.
This is the tool GitHub for Windows, which can be obtained here: https://git-scm.com/download/win
The problem is that we are under company proxy and there is no option to change it. Luckily there is a solution.
1. Open the file
C:\Users\YOURNAME\.gitconfig
2. Add there two lines
[http]
proxy = http://YOUR_COMPANY_PROXY:8080
[https]
proxy = http://YOUR_CONPANY_PROXY:8080
OK, all set!
By I feel using command line is more convenient.
1. clone
git clone REPOSITORY URL
for example
2. update
1) cd to the folder for example Spark
2) git pull
For the turorial of Git, we can check the great website, many thanks to the author!
http://rogerdudler.github.io/git-guide/
If you think this article is useful, please click the ads on this page to help. Thank you very much.
After install or unzip these components, configure them in the system variables, for example:
JAVA_HOME --> C:\Program Files\Java\jdk1.7.0_60
MAVEN_HOME --> D:\Program Files\Dev\apache-maven-3.2.3
SCALA_HOME --> C:\Program Files (x86)\scala
SBT_HOME --> C:\Program Files (x86)\sbt\
and append the following items in PATH %JAVA_HOME%\bin; %MAVEN_HOME%\bin;%SCALA_HOME%\bin;%SBT_HOME%\bin;
Note: it's mandatory to enable HTTP proxy if you are in a company firewall.
c) Configure Maven settings if under proxy
We can copy a settings.xml file from %MAVEN_HOME%\conf\settings.xml, and when it's under proxy, please copy it to C:\Users\*****\.m2 , and enable proxy
After modification:
3. Create project
Let's create a sample project called HelloSpark
1) File --> New --> Project..., please choose Scala
After click Finish, it will pop up a dialog, we can choose "New Window"
2). Add maven support
Right click project name and choose "Add Framework Support...", please scroll down and select "Maven"
Double click pom.xml and add the following content with existing content of pom.xml
After pasted the content, on the top right it will pop up a dialog, please choose Enable Auto-Import and maven will start downloading specified dependencies.
Or you can do it via
right click project name--> Maven --> Reimport
3) Create a folder for scala
expand project file structure, src--> main, right click main, New--> Directory,
name it as Scala
Then add this new folder "Scala" to project source
File--> Project Structure (shortcut Ctrl+Alt+Shift+S)
Modules--> scala -->Source , and as the screenshot shows, Click 1, 2 and 3, the result will display 4.
4) Create a scala class
Right click scala folder, new Scala class
Add modify the content as the screenshot.
Also please be noted
1) org.apache.spark.SparkContext need be imported.
2) create a file called pagecounts,
3) This program is to read the content from a file named pagecounts, and then print out the first 10 lines, and also print out the total line counts of this file.
You can put arbitrary content in pagecounts, a sample file can be viewed here. If you place in another folder, please modify the file path accordingly.
5) Add Spark jar file
We need to download and Spark latest package and unzip it
Go to: https://spark.apache.org/downloads.html
Downoad spark package, you can choose 2.4 or 2.6 based on your requirement. For example, a sampe spark-1.4.0-bin-hadoop2.4.tgz can be downloaded here.
After unzip, we can add the package in our project, click OK with the popup.
6) Set run configuration
in the IntelliJ menu, Run-->Edit Configuration, please choose Application and set up the content as the screenshot below
Final: Run it!
Click the run button on the toolbar, and the result is good!
Please note that in the beginning it will display SLF4J multiple binding problem and Winutil problem like java.io.IOException: Could not locate executable null\bin\winutils.exe in the Hadoop binaries. These can be ignored for now.
Happy Spark!
If you think this article is useful, please click the ads on this page to help. Thank you very much.