Output & Running
Hello, World
R has three ways to put text on screen and the choice matters here:
cat writes exactly what you give it, print adds the [1] index marker, and a bare expression auto-prints only in an interactive console.cat("Hello, World!\n")print "Hello, World!"That last point is why every R example on this page prints explicitly. A script run with
Rscript — and the WebR runtime behind the R column — does not auto-print a top-level value, so a bare x that shows a result at the console produces nothing at all. Nushell is the opposite: the last pipeline of a script prints itself, and only the ones before it need print.print, cat and the [1] marker
R's
print shows the value the way the console would, complete with the [1] that tells you where in the vector each displayed line starts. cat concatenates and prints raw, which is why it takes several arguments.numbers <- c(10, 20, 30)
print(numbers)
cat(numbers, "\n")
cat("total:", sum(numbers), "\n")let numbers = [10 20 30]
print $numbers
print ($numbers | str join " ")
print $"total: ($numbers | math sum)"The
[1] is a vector-indexing aid that becomes noise when you only wanted the numbers, which is exactly why R users reach for cat in scripts and print at the console. Nushell has one printer and it renders each type in its own way — a list as a bordered table, a record as key-value rows, a string as itself. That is why the first Nushell line here draws a table where the R line draws [1] 10 20 30: neither is decoration, both are the runtime showing you the shape of what you have.The Pipeline You Already Have
The pipe is the same idea
An R user does not need convincing that a pipeline over structured data is a good idea — that is what dplyr is, and it is why this page is shorter than it would be for a shell user. The vocabulary lines up almost word for word.
people <- data.frame(
name = c("ada", "grace", "alan"),
age = c(36, 45, 41),
city = c("london", "boston", "london")
)
result <- subset(people, city == "london")[, c("name", "age")]
print(result)let people = [
[name age city];
[ada 36 london]
[grace 45 boston]
[alan 41 london]
]
$people | where city == "london" | select name agefilter is where, select is select, mutate is insert or update, group_by plus summarize is group-by plus items, and arrange is sort-by. The base-R version is shown here rather than the dplyr one so that the comparison does not depend on a package, but the dplyr spelling is the closer match and the next rows use it. What Nushell adds is not a better pipeline; it is that this same pipeline reaches outside the process — over the filesystem, an HTTP response or another program's output — which is the subject of the last two sections.R's |> and Nushell's | differ in one way that matters
R gained a native pipe
|> in 4.1, and it works the way Nushell's does: the value on the left becomes the first argument of the call on the right. If you learned magrittr's %>% first, this is the same rule minus the . placeholder.numbers <- c(4, 1, 3, 2)
# R's native pipe sends into the FIRST argument:
result <- numbers |> sort() |> head(2)
print(result)
# Which is why a second argument is written normally:
print(numbers |> sort(decreasing = TRUE))let numbers = [4 1 3 2]
# Nushell's pipe also sends into the first position:
let result = $numbers | sort | first 2
print $result
# And a flag is written normally too:
print ($numbers | sort --reverse)The difference that matters is what a stage can be. In R, every stage is a function call, so a pipeline lives inside one R process and the things it can reach are the things loaded into that process. In Nushell every stage is a command, and a command may be a builtin, one you defined, or an external program, so the same syntax spans "sort this list" and "read every file in this directory". R's equivalent of the second is
system(), which hands back a character vector you must then parse — the one place where an R user meets the problem this page is about.Vectors Are Not the Default
Arithmetic is not vectorized
This is the change an R user feels first and hardest. In R, an operator applied to a vector applies elementwise, so
numbers * 2 is a whole vector. In Nushell a list is a container, and $numbers * 2 is an error.numbers <- c(1, 2, 3)
print(numbers * 2)
print(numbers + c(10, 20, 30))
print(sqrt(numbers))let numbers = [1 2 3]
print ($numbers | each { |number| $number * 2 })
print ($numbers | zip [10 20 30] | each { |pair| $pair.0 + $pair.1 })
print ($numbers | math sqrt)So every elementwise operation becomes an explicit
each, which is more typing and — unlike Python, where the answer is to install numpy — has no library that gives it back. Read that as a genuine cost of the language, not a hidden feature you have yet to find. Two things soften it: the math commands (math sum, math avg, math sqrt, math stddev) do work over a whole list, and a table column pulled out with get age is a list you can feed straight to them. What has no counterpart at all is R's recycling, where a short vector repeats to meet a long one — and given how many R bugs recycling has caused silently, losing it is not obviously a loss.There is no scalar in R, and there is one here
R has no scalar type:
5 is a numeric vector of length one, which is why length(5) is 1 and 5[1] is legal. Nushell has real scalars, and the distinction is visible everywhere once you notice it.x <- 5
print(length(x)) # 1 — it is a vector of length one
print(is.vector(x))
print(x[1]) # a vector can always be indexedlet x = 5
print ($x | describe) # int — a scalar, not a container
print ([$x] | length) # a list of one is something you build
print ([$x].0)That single fact explains most of the differences on this page. Because everything in R is a vector, an operator can be elementwise, a function can accept any length, and
if historically took the first element of a longer condition — a source of bugs so persistent that R 4.2 made it an error. Because a Nushell value is either a scalar or a container, an operation on the wrong one is caught rather than broadcast, and length on an integer is an error rather than 1. Neither model is better in the abstract; the R one is built for data analysis and the Nushell one for a shell where a filename is a filename.Variables & Assignment
Assignment, and immutability by default
R's
<- rebinds freely, and = does the same thing outside a function call. Nushell's let binds once: changing a value needs mut, and reading one needs a $ that is absent from the declaration.name <- "Ada"
count <- 0
count <- count + 1
cat(name, count, "\n")let name = "Ada"
mut count = 0
$count += 1
print $"($name) ($count)"The
$ asymmetry catches everybody: it is let name = … to bind and $name to read, because the sigil belongs to the reference rather than the name. As for mut, you will need it less than you expect, because the accumulate-in-a-loop pattern is a pipeline here. Note also the rule with no R analogue: a mut variable cannot be captured by a closure, so each { |item| $total += $item } is rejected — which pushes you to reduce, and which exists because a closure may one day run in parallel.Copy-on-modify, made into a rule
R's copy-on-modify semantics mean this already behaves the way an R user expects, which makes it the friendliest row on the page — and the one that hides the difference, so read the note.
original <- c(1, 2, 3)
copy <- original
copy[1] <- 99
print(original) # unchanged — R copied on modify
print(copy)let original = [1 2 3]
let copy = $original | update 0 99
print $original # unchanged — there is no modify to copy on
print $copyR gives you the appearance of value semantics through an optimization: the vector is shared until somebody writes to it, and then it is copied. That is why a large data frame can quietly double in memory at an assignment, and why
tracemem exists. Environments and R5/R6 reference classes escape it entirely and do alias. In Nushell there is no modification operation at all, so nothing is copied on write because nothing is ever written — update returns a new list, and the old one and the new one may share whatever the runtime finds convenient with no observable difference.Indexing — Every Rule Differs
Indexing starts at zero
R is one of the few languages that indexes from 1, and Nushell is not one of them. This is the most mechanical difference on the page and the one most likely to produce a wrong answer rather than an error.
friends <- c("Grace", "Alan", "Edsger")
print(friends[1]) # the FIRST element
print(friends[3])
print(friends[length(friends)])let friends = [Grace Alan Edsger]
print $friends.0 # the FIRST element
print $friends.2
print ($friends | last)Off-by-one is the least of it — see the next row for the part that actually bites. Two smaller notes: there is no
length(x)-based way to reach the end here, because last exists and reads better; and indexing past the end gives an error in Nushell rather than R's NA, which is one place the stricter language saves you from a silently wrong result.Negative indexing means the opposite thing
R's negative index excludes:
numbers[-1] is everything but the first element. Almost every other language reads it as "count from the end", so this is the R idiom most likely to be mistranslated in either direction.numbers <- c(10, 20, 30, 40)
print(numbers[-1]) # DROPS the first element
print(numbers[c(-1, -2)])
print(numbers[2:3]) # inclusive range, 1-basedlet numbers = [10 20 30 40]
print ($numbers | skip 1) # drop the first element
print ($numbers | skip 2)
print ($numbers | slice 1..2) # inclusive range, 0-basedNushell has no negative indexing at all, which removes the ambiguity rather than resolving it —
skip and drop say which end and how many, and last gets the final element. The range rule is worth pinning down because it is the one thing that does match: slice's range is inclusive at both ends, exactly like R's 2:3, so only the zero base moves. That makes numbers[2:3] into slice 1..2, which is a subtraction rather than the off-by-one-at-one-end conversion that Python would demand.Named access: $ becomes .
R's list is a record with names, reached with
$ or [[ ]]. Nushell's record is the same idea reached with a dot, and get is the pipeline form for when the key is in a variable.person <- list(name = "Ada", age = 36)
print(person$name)
print(person[["age"]])
print(names(person))let person = {name: "Ada", age: 36}
print $person.name
print ($person | get age)
print ($person | columns)One R behavior to leave behind:
$ does partial matching, so person$na returns the name field, and a typo can silently resolve to the wrong element. [[ ]] does not, which is why careful R code prefers it. Nushell has no partial matching anywhere — $person.na is a Name not found error that suggests name. Note also that names() becomes columns, because a record and a single table row are the same thing here.Data Frames Become Tables
A data frame becomes a table
A Nushell table is a list of records that share keys, which is a data frame described from the other direction — R builds it from columns, Nushell writes it as rows.
people <- data.frame(
name = c("ada", "grace"),
age = c(36, 45)
)
print(people)
print(nrow(people))
print(ncol(people))let people = [
[name age];
[ada 36]
[grace 45]
]
print $people
print ($people | length)
print ($people | columns | length)That difference in construction is real. A data frame is fundamentally column-major: each column is a vector of one type, which is what makes
people$age free and vectorized arithmetic over it natural. A Nushell table is row-major, so get age gathers a column by walking the rows, and there is no promise that every row has the same columns — a table with a missing cell is legal and shows as empty. For analysis work that is a downgrade; for a shell, where the rows come from a directory listing or a JSON array, it is the honest shape.Getting a column out
Pulling a column out gives you a list, and from there the
math commands work over it the way R's vectorized functions do. This is the shape in which an R user gets vectorization back.people <- data.frame(
name = c("ada", "grace", "alan"),
age = c(36, 45, 41)
)
print(people$age)
print(mean(people$age))
print(people[people$age > 40, "name"])let people = [
[name age];
[ada 36]
[grace 45]
[alan 41]
]
print ($people | get age)
print ($people | get age | math avg)
print ($people | where age > 40 | get name)The third line is where the two diverge in readability. R's
people[people$age > 40, "name"] repeats the data frame name inside its own subscript, which is why dplyr exists and why filter(people, age > 40) reads so much better. Nushell's where age > 40 is closer to the dplyr spelling than to the base-R one — the column names are in scope inside the condition, with no $ and no repetition. Note that get on a missing column errors rather than returning NULL, which is the general pattern: this language would rather stop than hand you nothing.dplyr Verbs, Command by Command
mutate becomes insert
R adds a column by assigning to a name that does not exist yet, and dplyr's
mutate wraps that in a pipeline-friendly verb. Nushell splits the job into three commands by what already exists.orders <- data.frame(
item = c("book", "pen"),
price = c(12, 3),
quantity = c(2, 5)
)
orders$total <- orders$price * orders$quantity
print(orders)let orders = [
[item price quantity];
[book 12 2]
[pen 3 5]
]
$orders | insert total { |row| $row.price * $row.quantity }insert adds a column and fails if it is already there, update changes one and fails if it is not, and upsert does whichever applies. R's orders$total <- … is all three at once, which is convenient right up to the typo that silently creates a column named totl beside the one you meant. The closure receives the whole row, so any column is available — and note the vectorization difference: R computes price * quantity as one vector operation over the whole column, while Nushell runs the closure once per row.arrange becomes sort-by
Base R sorts a data frame by computing an ordering vector with
order() and using it as a row subscript — the indirection that arrange() exists to hide. Nushell names the columns directly.people <- data.frame(
name = c("grace", "ada", "alan"),
age = c(45, 36, 41),
city = c("boston", "london", "london")
)
print(people[order(people$age), "name"])
print(people[order(-people$age), "name"])
print(people[order(people$city, people$age), "name"])let people = [
[name age city];
[grace 45 boston]
[ada 36 london]
[alan 41 london]
]
print ($people | sort-by age | get name)
print ($people | sort-by age --reverse | get name)
print ($people | sort-by city age | get name)The
-people$age trick for descending order is a small piece of R folklore worth naming: it works because negating the values reverses their order, which means it does not work on a character column at all — you need order(people$name, decreasing = TRUE) there. Nushell's --reverse is a flag on the command, so it applies to any type. Multiple sort keys work the same way in both, listed in priority order.select, and dropping columns
Keeping columns is
select in both vocabularies. Dropping one is where base R gets awkward — you compute the complement of the names yourself — and where Nushell has a word for it.people <- data.frame(
name = c("ada", "grace"),
age = c(36, 45),
secret = c("x", "y")
)
print(people[, c("name", "age")])
print(people[, setdiff(names(people), "secret")])let people = [
[name age secret];
[ada 36 x]
[grace 45 y]
]
print ($people | select name age)
print ($people | reject secret)dplyr closes that gap with
select(people, -secret), which is the fair comparison and reads well; the base-R version is shown to make the point that a verb has to exist somewhere. reject is the exact counterpart and, like select, takes several names. Both error on a missing column; reject --optional is the lenient form, for when the only goal is that the column not be there.Grouping & Summarizing
group_by and summarize
Base R spells this with a formula —
total ~ customer reads "total broken down by customer" — which is compact and unlike anything else in either language. Nushell splits it into grouping and then summarizing.orders <- data.frame(
customer = c("ada", "grace", "ada"),
total = c(30, 45, 12)
)
result <- aggregate(total ~ customer, data = orders, FUN = sum)
print(result)let orders = [
[customer total];
[ada 30]
[grace 45]
[ada 12]
]
$orders
| group-by customer
| items { |customer, rows| {customer: $customer, total: ($rows.total | math sum)} }R wins on brevity here and it is worth saying so; the formula interface is one of the language's genuinely good ideas, and dplyr's
group_by(customer) |> summarize(total = sum(total)) is barely longer. What Nushell gives back is that the intermediate is an ordinary value you can look at: group-by customer alone returns a record whose keys are the customers and whose values are whole tables, so you can inspect it, filter it, or summarize several ways without recomputing. items is what walks a record giving you each key and value together, and $rows.total pulls one column from a whole table.table() becomes uniq --count
R's
table() is the fastest way in any language to count occurrences, and it returns a named vector that prints as a little labelled row. Nushell returns a two-column table.fruits <- c("apple", "pear", "apple", "fig", "pear", "apple")
counts <- table(fruits)
print(counts)
print(sort(counts, decreasing = TRUE))let fruits = [apple pear apple fig pear apple]
print ($fruits | uniq --count)
print ($fruits | uniq --count | sort-by count --reverse)The difference in return type is the whole comparison. R's
table gives you an object with its own printing, its own names(), and its own quirks — sorting it keeps the names attached, but turning it into a data frame needs as.data.frame and produces columns called Var1 and Freq. Nushell's uniq --count returns an ordinary table with value and count columns, so every table command already applies and nothing needs converting. Plain uniq, with no flag, gives just the distinct values, which is R's unique().The apply Family
sapply and lapply become each
The apply family collapses to one command.
sapply and lapply are both each, Filter is where, and Reduce is reduce.numbers <- c(1, 2, 3, 4)
print(sapply(numbers, function(n) n * n))
print(Filter(function(n) n %% 2 == 0, numbers))
print(Reduce(function(a, b) a + b, numbers, 0))let numbers = [1 2 3 4]
print ($numbers | each { |number| $number * $number })
print ($numbers | where { |number| $number mod 2 == 0 })
print ($numbers | reduce --fold 0 { |number, total| $total + $number })What disappears with
sapply is its most notorious property: it simplifies the result, returning a vector when it can and a list when it cannot, so the type of what you get back depends on the data rather than on the code. That is why vapply exists and why tidyverse code uses the map_* family with the type in the name. each over a list always returns a list. Note reduce's argument order, which is the reverse of R's: the element comes first and the accumulator second, so it is { |number, total| … }. --fold supplies the initial value, and without it the first element becomes it.apply over rows
R's
apply with MARGIN = 1 walks the rows of a matrix or data frame. This is the operation where Nushell's row-major model is the natural one and R's column-major model is fighting you.measurements <- data.frame(
width = c(2, 3),
height = c(4, 5)
)
areas <- apply(measurements, 1, function(row) row["width"] * row["height"])
print(areas)let measurements = [
[width height];
[2 4]
[3 5]
]
print ($measurements | each { |row| $row.width * $row.height })The R version has a trap that has bitten every R user at least once:
apply converts the data frame to a matrix first, and a matrix has one type, so a data frame with any character column turns every number into a string and the multiplication fails or lies. That is why rowwise() and pmap exist in the tidyverse. Nushell's each over a table hands you a real record per row with every value keeping its own type, because a table was always a list of records. This row is the clearest case on the page where the shell's model is the better fit.NA Has No Counterpart
NA becomes null, and it propagates less
R's
NA is missingness as a first-class idea: it propagates through arithmetic, it is contagious in comparisons, and there is a separate NA for each type. Nushell has null, which is absence rather than missingness.values <- c(1, 2, NA, 4)
print(sum(values)) # NA — it spreads
print(sum(values, na.rm = TRUE)) # 7
print(NA > 1) # NA, not TRUE or FALSElet values = [1 2 null 4]
print ($values | compact | math sum) # 7 — compact drops the nulls first
print ($values | compact | length) # 3
print (null == 1) # false — a real boolean, not nullThe three lines show three different answers to the same question. R's
sum(values) returns NA because it genuinely does not know the total, and na.rm = TRUE is you stating that you accept an answer computed from what is there. Nushell's math sum does neither — it refuses the list outright with only list<number> … is supported, because [1 2 null 4] is not a list of numbers. So you must call compact, which is na.rm = TRUE spelled as a pipeline stage. The third line is the one to remember: NA > 1 is NA in R, so a comparison involving missing data is itself unknown, while null == 1 is plainly false. For statistics that contagion is the point and losing it is a real loss; for a shell, where a null is a directory entry with no size rather than an unrecorded measurement, plain false is right. If missingness matters in your data, count the nulls yourself, because nothing here will remind you.Strings
The string toolkit
R's string functions are base functions taking the subject first; Nushell's are subcommands of
str taking it from the pipeline. The names differ but the jobs line up one to one.text <- " Hello, World "
trimmed <- trimws(text)
print(toupper(trimmed))
print(nchar(trimmed))
print(substr(trimmed, 1, 5))
print(gsub("World", "Nushell", trimmed))let text = " Hello, World "
let trimmed = $text | str trim
print ($trimmed | str uppercase)
print ($trimmed | str length)
print ($trimmed | str substring 0..4)
print ($trimmed | str replace "World" "Nushell")The substring row is the one to write down. R's
substr(x, 1, 5) is 1-based and inclusive; Nushell's str substring 0..4 is 0-based and also inclusive, so the conversion is "subtract one from both ends" rather than the mixed adjustment other languages need. Two Nushell details: str upcase and str downcase were deprecated in 0.114 in favor of str uppercase and str lowercase, and str replace replaces the first match by default — R's gsub is global and its sub is not, so gsub maps to str replace --all.paste, sprintf and splitting
R has
paste, paste0 and sprintf for building strings; Nushell has one interpolated-string form, $"…($expression)…", whose holes take any expression including a whole pipeline.name <- "Ada"
age <- 36
print(paste0("Name: ", name, ", age: ", age))
print(sprintf("%s is %d", name, age))
print(strsplit("a,b,c", ",")[[1]])let name = "Ada"
let age = 36
print $"Name: ($name), age: ($age)"
print $"($name) is ($age)"
print ("a,b,c" | split row ",")The R detail worth remembering here is that
paste is vectorized: given vectors it returns a vector of results, which is enormously useful and also why paste with a stray length-3 vector silently produces three strings where you wanted one. And strsplit returns a list of character vectors even for a single input, which is why the [[1]] is there — a wart the tidyverse's str_split_1 was added to fix. Nushell's split row returns the list directly, and there is no vectorization to be surprised by.Regular expressions
Both have regular expressions. R's live in
grep, grepl, sub and gsub, with a value = TRUE flag deciding whether you get matches or positions; Nushell uses =~ for matching and parse for extraction.lines <- c("ERROR disk full", "INFO started", "ERROR timeout")
print(grep("^ERROR", lines, value = TRUE))
print(length(grep("^ERROR", lines)))
print(sub("^([A-Z]+) (.*)$", "\\1/\\2", lines[1]))let lines = ["ERROR disk full" "INFO started" "ERROR timeout"]
print ($lines | where $it =~ '^ERROR')
print ($lines | where $it =~ '^ERROR' | length)
print ($lines.0 | parse --regex '^(?<level>[A-Z]+) (?<message>.*)$')Two R details that cost people time:
grep returns indices by default and the matched values only with value = TRUE, and backreferences in a replacement are written \\1 — doubled, because R string literals process backslashes before the regex engine sees them. Nushell's parse takes a different approach to extraction entirely: it returns a table with a column per named group, so running it over a whole list gives you a table of everything, ready to filter and sort. The single quotes matter — a single-quoted Nushell string passes backslashes through untouched, which is what a regex wants.Functions Become Commands
function becomes def, with types
Both return the last expression with no
return needed, and both support default arguments. What Nushell adds is types in the signature; what it takes away is R's named-argument calling.greet <- function(name, greeting = "Hello") {
paste0(greeting, ", ", name, "!")
}
print(greet("Ada"))
print(greet("Ada", "Welcome"))
print(greet(greeting = "Hi", name = "Ada"))def greet [name: string, greeting: string = "Hello"] {
$"($greeting), ($name)!"
}
print (greet "Ada")
print (greet "Ada" "Welcome")
print (greet --help | lines | where $it =~ "greet " | first)That third R line is the loss. R lets any argument be passed by name in any order, with partial matching on the name — a genuinely convenient feature that also means
greet(gr = "Hi") works and greet(g = "Hi") is ambiguous only if another parameter starts with g. Nushell's positional parameters are positional; anything you want to pass by name is declared as a --flag, which is the next row. In exchange, the signature is real documentation: the third Nushell line pulls it out of greet --help, which prints > greet <name> (greeting) — required parameters in angle brackets, optional ones in parentheses, generated from the declaration and impossible to leave stale. R needs a separate .Rd file or roxygen comments kept in step by hand.Named arguments become flags
Anything you would pass to an R function by name becomes a flag here, declared with a leading
--. A flag with no type is a switch: present means true, absent means false.report <- function(rows, separator = ", ", loud = FALSE) {
text <- paste(rows, collapse = separator)
if (loud) toupper(text) else text
}
print(report(c("ada", "grace")))
print(report(c("ada", "grace"), separator = " | ", loud = TRUE))def report [rows: list, --separator: string = ", ", --loud] {
let text = $rows | str join $separator
if $loud { $text | str uppercase } else { $text }
}
print (report [ada grace])
print (report [ada grace] --separator " | " --loud)So
loud = FALSE as a default disappears — --loud needs no default because absence is the default. The larger point is that the command now behaves exactly like a builtin: it takes flags, it appears in help, it tab-completes at the prompt, and a caller cannot tell it was not shipped with Nushell. An R function with named arguments is just an R function; there is no sense in which report becomes part of the language's command vocabulary, because R has no command vocabulary to join.A command that reads from the pipeline
R's pipe fills the first parameter, so a function is pipeable by having its data argument first — a convention the tidyverse follows rigorously. Nushell has a separate channel:
$in is the pipeline input, distinct from the parameters.# R functions take arguments; the pipe just fills the first one.
shout <- function(words) {
toupper(words)
}
print(c("ada", "grace") |> shout())def shout []: list -> list {
$in | each { |word| $word | str uppercase }
}
print ([ada grace] | shout)The distinction sounds academic until you want both. A Nushell command can take pipeline input and positional arguments and flags, each in its own place, so
$data | my-report --format json has nowhere for the three to collide. In R everything is a parameter, so a function that is pipeable and also configurable relies on the first parameter being the data by convention, and nothing enforces it — which is exactly why base R's inconsistency here (sapply(X, FUN) against Map(f, ...)) makes some functions awkward to pipe into. The : list -> list annotation declares the input and output types and is optional.Files & the Working Directory
Reading a file
R's
readLines and writeLines take a filename and a character vector, which is close enough to open and save that this row is mostly here to set up the next two.writeLines(c("first line", "second line"), "rnushellread-notes.txt")
lines <- readLines("rnushellread-notes.txt")
cat(length(lines), "lines\n")
cat(lines[1], "\n")"first line\nsecond line\n" | save --force rnushellread-notes.txt
let lines = open rnushellread-notes.txt | lines
print $"($lines | length) lines"
print $lines.0Both columns write the file before reading it, and both use a filename prefix no other row on this page uses. That is not tidiness: in the browser every run starts from a fresh engine and an empty filesystem, so an example must create whatever it reads, and locally the examples share a real temporary directory where one row could otherwise pick up another's leftovers. Two small Nushell notes:
lines drops the trailing newline so there is no chomp equivalent to forget, and save refuses to overwrite without --force, which is the opposite of R's silently-clobbering writeLines.Listing files, and what ls returns
This is where the two genuinely part company. R's
list.files returns a character vector of names, so every further question about a file — its size, its modification time, whether it is a directory — is a separate call taking that name back to the filesystem.dir.create("rnushellglob-dir", showWarnings = FALSE)
for (name in c("alpha.txt", "beta.txt", "gamma.log")) {
cat("xxxxxxxxxx", file = file.path("rnushellglob-dir", name))
}
found <- list.files("rnushellglob-dir", pattern = "\\.txt$")
cat(length(found), "text files\n")
cat(file.size(file.path("rnushellglob-dir", found[1])), "bytes in the first\n")mkdir rnushellglob-dir
for name in [alpha.txt beta.txt gamma.log] {
"xxxxxxxxxx" | save --force $"rnushellglob-dir/($name)"
}
let found = ls rnushellglob-dir/*.txt
print $"($found | length) text files"
print $"($found.0.size) in the first"Nushell's
ls returns a table with name, type, size and modified already filled in, so "the three biggest files modified this week" is ls | where modified > (date now) - 7day | sort-by size --reverse | first 3 with no second trip to disk. R can get there with file.info(), which does return a data frame — but you have to know to ask, and you have to pass it the names you already got. The size also prints as 10 B rather than 10, because it is a filesize value that compares correctly against 1mb and formats itself when shown. This is the concrete form of the claim that the pipeline reaches outside the process.Calling another program hands back text
R's
system(..., intern = TRUE) is how an analysis reaches a command-line tool, and what comes back is a character vector. Nushell has the same capability spelled ^command, with the same consequence.# R's escape hatch returns a character vector — text to parse:
output <- system("echo one two three", intern = TRUE)
fields <- strsplit(output, " ")[[1]]
cat(length(fields), "fields\n")
print(fields)# Nushell CAN call externals with ^, but this build has no processes.
# The point is what happens when you do: you get TEXT back.
let output = "one two three"
let fields = $output | split row " "
print $"($fields | length) fields"
print $fieldsThat consequence is the honest limit of everything above: the moment you call an external program you are back to parsing text, in either language. Nushell's answer is to ship builtin equivalents of the tools you would otherwise shell out to —
ls, du, ps, http get — so the escape hatch is needed less often, and to provide from ssv, from json and friends for when it is. R's answer is a package per tool. The in-browser build on this page has no processes at all, which is why the second column demonstrates the parsing rather than the call.CSV, JSON & the Format Zoo
Reading CSV
This is R's home ground:
read.csv is one call, it infers types per column, and the result is a data frame ready to work with. Nushell's from csv does the same job with the same one call.csv <- "name,age\nada,36\ngrace,45\n"
people <- read.csv(text = csv, stringsAsFactors = FALSE)
print(people)
print(class(people$age))let csv = "name,age\nada,36\ngrace,45\n"
let people = $csv | from csv
print $people
print ($people | describe)Two pieces of R history are visible in that call.
stringsAsFactors = FALSE had to be written on every read.csv until R 4.0 changed the default, because character columns were silently converted to factors and comparing one against a string then failed in ways that took an afternoon to find. And read.csv guesses types by scanning, which is why a column of identifiers with a leading zero arrives as an integer with the zero gone. Nushell infers types too and has the same class of hazard; describe is how you check what you actually got, and it is worth the habit in both languages.JSON needs a package in R and nothing here
Base R ships no JSON parser.
jsonlite is the near-universal answer and is one install.packages away, but it is a dependency, and this row shows what a script without it is left doing.# Base R has no JSON reader; jsonlite is the standard answer.
# Without it, this is what you are reduced to:
text <- '{"port": 8080, "debug": true}'
port <- sub('.*"port"[^0-9]*([0-9]+).*', "\\1", text)
cat("port:", port, "\n")let text = '{"port": 8080, "debug": true}'
let config = $text | from json
print $"port: ($config.port)"
print $"type: ($config | describe)"The regex in the R column is not a straw man — it is what people actually write when they need one field out of a small JSON blob and do not want a dependency for it, and it breaks the moment the key order changes or a nested object contains the same key name. With
jsonlite, fromJSON(text)$port is one line and the comparison becomes fair. What remains is that Nushell has no dependency to install and no import to write, and that the same from family covers YAML, TOML, XML and INI with identical syntax — so a script that reads a config file does not care which format the config is in.Converting between formats
Once a format is a table, converting to another is one more command. This row is one line on the Nushell side and that is the whole content of it.
# Needs jsonlite for the second half; base R can only get you here.
csv <- "name,age\nada,36\ngrace,45\n"
people <- read.csv(text = csv, stringsAsFactors = FALSE)
rows <- apply(people, 1, function(row) {
paste0('{"name":"', row["name"], '","age":', row["age"], '}')
})
cat("[", paste(rows, collapse = ","), "]\n", sep = "")let csv = "name,age\nada,36\ngrace,45\n"
$csv | from csv | to json --rawThe R column shows the hand-rolled version, and note what
apply did to it: converting the data frame to a matrix turned the ages into strings, which is why they have to be pasted in unquoted by hand and why a value containing a quote would produce invalid JSON. jsonlite::toJSON(people) is the real answer and is one line. The durable difference is not the line count but the coverage: from X | to Y works for every pair Nushell knows, so CSV to JSON, TOML to YAML and XML to CSV are the same command with different words, where R needs a package per format and each has its own conventions about what a row is.⚠ Gotchas for R Users
A condition must be a real boolean
R treats
0 as false and any other number as true, and its logicals are numbers underneath — TRUE + TRUE is 2, which is what makes sum(x > 5) the idiomatic way to count matches. Nushell has no coercion at all.count <- 0
if (count > 0) cat("has items\n") else cat("empty\n")
# R coerces 0 and 1 to FALSE and TRUE:
if (1) cat("one is true\n")
# Counting matches, the idiom that relies on logicals being numbers:
print(sum(c(TRUE, TRUE, FALSE)))let count = 0
print (if $count > 0 { "has items" } else { "empty" })
# Nushell does not coerce: an int is not a condition.
print (1 | into bool)
# Counting matches, which is what sum(x > 5) does in R:
print ([true true false] | where {|value| $value} | length)So
if $count { … } is not a shortcut, it is an error — can't convert int to boolean, at the moment the line runs. And sum(x > 5), which counts matches by relying on logicals being numbers, becomes $x | where {|value| $value > 5} | length — the last line of each column shows the same count reached the two different ways. into bool exists for deliberate conversion and is narrower than you would guess: it takes a number, a string holding a number, or the strings "true" and "false", and errors on anything else, so "0" | into bool is false while "yes" | into bool is an error. Losing R's coercion costs brevity and removes the class of bug where a legitimate zero silently reads as "nothing here".Arguments are separated by spaces
Nushell is a shell first, so calling a command is shell syntax: the name, then arguments separated by spaces. Commas in a list literal are optional and usually left out.
add <- function(left, right) left + right
print(add(2, 3))
print(c(1, 2, 3)) # commas everywheredef add [left: int, right: int] { $left + $right }
print (add 2 3) # spaces, not commas
print [1 2 3] # a list literal has no commas eitherThe failure mode is quiet, which is why this is a gotcha rather than a note.
add 2, 3 is not a syntax error — it passes the string "2," and then 3, which errors on the type here but would do something wrong and silent in a command taking strings. The other half of the rule is that a command call used as an expression needs parentheses: print (add 2 3), because print add 2 3 means printing four separate things. Once both halves are internalized the syntax stops surprising you, and until then it is the most common source of confusion.par-each exists, but not in this build
Nushell has
par-each, which runs a closure across threads and is one of the reasons people reach for it. The in-browser build behind this page has no threads, so it is not available here.numbers <- c(1, 2, 3, 4, 5)
# R's parallelism is a package away: parallel::mclapply, future, foreach.
squares <- sapply(numbers, function(n) n * n)
print(squares)let numbers = [1 2 3 4 5]
# par-each is real Nushell, but needs threads this build does not have.
let squares = $numbers | each { |number| $number * $number }
print ($squares | str join ",")each does the same work in the same order, so nothing on this page is wrong — only slower on a large list than the same script would be at a terminal. It is worth knowing what the real thing costs elsewhere: because Nushell closures capture by value and nothing is shared, par-each needs no lock and no cluster setup, which is a much smaller promise than R's parallel package makes with its fork-or-socket-cluster choice and its rules about what can be sent to a worker. Result order is not guaranteed with par-each; pipe through sort-by when it matters.Nothing survives between runs
This is a property of the page rather than of the languages, and it changes how every file example above should be read. Each click here starts a brand-new Nushell engine with an empty filesystem.
# An R session accumulates: the workspace persists across commands,
# which is why rm(list = ls()) is a reflex and .Rdata surprises people.
marker <- "written by this run"
writeLines(marker, "rnushellleftover-marker.txt")
cat(file.exists("rnushellleftover-marker.txt"), "\n")# In the browser each run gets a FRESH engine and an EMPTY filesystem.
"written by this run\n" | save --force rnushellleftover-marker.txt
print ("rnushellleftover-marker.txt" | path exists)So an example must create everything it reads, and nothing one row leaves behind is visible to another. That is the opposite of the R habit, where the workspace persists across an entire session and a variable defined an hour ago is still there — the reason
rm(list = ls()) is a reflex before running a script, and the reason a script that works interactively can fail from a clean session. Treat every row on this page the way you should treat an R script: as if nothing exists that it did not create.